mediumMultiple Choice
PDE Practice Question: A retail company processes real-time clickstream…
A retail company processes real-time clickstream data using Cloud Pub/Sub and Dataflow. The pipeline aggregates events by user session and writes to Bigtable for low-latency queries. However, users report that session data is sometimes missing or duplicated. What is the most likely cause?
⚠ Common exam trap
Google Cloud often tests the misconception that Pub/Sub provides exactly-once delivery or that Dataflow automatically deduplicates messages from Pub/Sub, when in fact Pub/Sub is at-least-once and Dataflow requires explicit deduplication for idempotent processing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pub/Sub provides at-least-once delivery, and Dataflow does not deduplicate by default.
D is correct because Pub/Sub offers at-least-once delivery, meaning the same message may be delivered multiple times. Dataflow does not automatically deduplicate messages unless explicitly configured (e.g., using idempotent sinks or custom deduplication logic). Without deduplication, the same session event can be processed more than once, leading to duplicate session data in Bigtable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Session windowing is configured with too short a gap duration.
Why it's wrong here
A short gap duration splits one session into several windows, which can duplicate aggregates but does not drop events, so missing data remains unexplained. It is tempting because gap duration governs session boundaries, yet it would be correct only when sessions are being fragmented, not when records vanish.
- ✗
Bigtable schema design causes row key collisions.
Why it's wrong here
Bigtable row key collisions overwrite or merge distinct rows, producing missing data, but they cannot create duplicated session records. It is tempting because row key design genuinely affects Bigtable performance, yet it would be the cause only when keys are non-unique, not when duplicates appear.
- ✗
Dataflow's default behavior discards late-arriving data.
Why it's wrong here
Dataflow's default windowing does not silently discard late data; late elements are dropped only when allowed lateness is exceeded or the accumulation mode is set to discarding. It is tempting because late events are a real streaming concern, but it would be correct only with explicit lateness or accumulation settings.
- ✓
Pub/Sub provides at-least-once delivery, and Dataflow does not deduplicate by default.
Why this is correct
Pub/Sub guarantees at-least-once delivery, so messages can be redelivered after acknowledgement timeouts or retries. Dataflow's default pipeline performs no deduplication, so these redeliveries produce duplicate session records, while late or dropped acknowledgements also cause apparent gaps.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.