easyMultiple Choice
PDE Practice Question: Designing a streaming data pipeline to process…
A company is designing a streaming data pipeline to process real-time clickstream events. They need to aggregate events by session window with a 5-minute gap and enable exactly-once processing semantics. Which Google Cloud service should they use?
⚠ Common exam trap
Google Cloud often tests the distinction between stateless serverless services (like Cloud Functions) and stateful stream processing engines (like Dataflow), leading candidates to incorrectly choose Cloud Pub/Sub with Cloud Functions because they overlook the need for session window state management and exactly-once semantics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataflow with Apache Beam
Cloud Dataflow with Apache Beam is the correct choice because it provides native support for session windows with a 5-minute gap duration and exactly-once processing semantics via its sink and source integrations. Dataflow's Beam SDK allows you to define session windows using `Window.into(Sessions.withGapDuration(Duration.standardMinutes(5)))`, and its checkpointing and idempotent writes ensure exactly-once delivery even in failure scenarios.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Pub/Sub with Cloud Functions
Why it's wrong here
Cloud Functions offers no session-window aggregation with a five-minute gap and no exactly-once guarantees across retries. It is tempting because Pub/Sub plus Functions is a common lightweight event pipeline, and would be correct for simple per-event processing where at-least-once delivery and stateless handling suffice.
- ✓
Cloud Dataflow with Apache Beam
Why this is correct
Cloud Dataflow with Apache Beam satisfies both constraints: Beam's session windows natively aggregate events separated by a 5-minute gap, and Dataflow's exactly-once processing guarantees deduplicate records end-to-end. Pub/Sub alone lacks windowing, while Dataproc requires manual checkpointing, so neither meets the stated semantics.
- ✗
Cloud Dataproc with Spark Streaming
Why it's wrong here
Spark Streaming's micro-batch model gives at-least-once semantics by default, and Dataproc requires cluster management for a continuous pipeline. It is tempting because Spark supports session windows, and would be correct for batch-oriented analytics on existing Hadoop workloads rather than a managed exactly-once streaming service.
- ✗
Cloud Bigtable with Dataflow templates
Why it's wrong here
Bigtable stores and serves data but performs no windowing; Dataflow templates cannot express session windows with a five-minute gap or exactly-once state. It is tempting because Bigtable suits high-throughput clickstream storage, and would be correct as the serving layer beneath a pipeline rather than the processing engine itself.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.