Courseiva
easyMultiple Choice

PDE Practice Question: Designing a streaming data pipeline to process…

A company is designing a streaming data pipeline to process real-time clickstream events. They need to aggregate events by session window with a 5-minute gap and enable exactly-once processing semantics. Which Google Cloud service should they use?

⚠ Common exam trap

Google Cloud often tests the distinction between stateless serverless services (like Cloud Functions) and stateful stream processing engines (like Dataflow), leading candidates to incorrectly choose Cloud Pub/Sub with Cloud Functions because they overlook the need for session window state management and exactly-once semantics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataflow with Apache Beam

Cloud Dataflow with Apache Beam is the correct choice because it provides native support for session windows with a 5-minute gap duration and exactly-once processing semantics via its sink and source integrations. Dataflow's Beam SDK allows you to define session windows using `Window.into(Sessions.withGapDuration(Duration.standardMinutes(5)))`, and its checkpointing and idempotent writes ensure exactly-once delivery even in failure scenarios.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Pub/Sub with Cloud Functions

    Why it's wrong here

    Cloud Functions offers no session-window aggregation with a five-minute gap and no exactly-once guarantees across retries. It is tempting because Pub/Sub plus Functions is a common lightweight event pipeline, and would be correct for simple per-event processing where at-least-once delivery and stateless handling suffice.

  • ✓

    Cloud Dataflow with Apache Beam

    Why this is correct

    Cloud Dataflow with Apache Beam satisfies both constraints: Beam's session windows natively aggregate events separated by a 5-minute gap, and Dataflow's exactly-once processing guarantees deduplicate records end-to-end. Pub/Sub alone lacks windowing, while Dataproc requires manual checkpointing, so neither meets the stated semantics.

  • ✗

    Cloud Dataproc with Spark Streaming

    Why it's wrong here

    Spark Streaming's micro-batch model gives at-least-once semantics by default, and Dataproc requires cluster management for a continuous pipeline. It is tempting because Spark supports session windows, and would be correct for batch-oriented analytics on existing Hadoop workloads rather than a managed exactly-once streaming service.

  • ✗

    Cloud Bigtable with Dataflow templates

    Why it's wrong here

    Bigtable stores and serves data but performs no windowing; Dataflow templates cannot express session windows with a five-minute gap or exactly-once state. It is tempting because Bigtable suits high-throughput clickstream storage, and would be correct as the serving layer beneath a pipeline rather than the processing engine itself.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.