mediumMultiple Choice
PDE Practice Question: Building a real-time streaming pipeline using…
A company is building a real-time streaming pipeline using Pub/Sub and Dataflow to process clickstream data. The pipeline writes aggregated metrics to BigQuery every 10 seconds using a fixed window. During peak traffic, some windows produce duplicate rows in BigQuery. What is the most likely cause?
⚠ Common exam trap
Test-takers frequently confuse trigger behavior (Option B) with the root cause of duplicates, not realizing that duplicates stem from retry semantics in the sink, not from windowing or parallelism.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataflow is retrying BigQuery streaming inserts after a timeout, and the retries succeed even though the original insert succeeded.
Dataflow uses at-least-once semantics for streaming inserts into BigQuery. When a streaming insert times out, Dataflow retries the insert, and if the original insert actually succeeded but the acknowledgment was lost, the retry produces a duplicate row. This is a known behavior of BigQuery streaming inserts with retry logic.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Dataflow is retrying BigQuery streaming inserts after a timeout, and the retries succeed even though the original insert succeeded.
Why this is correct
BigQuery streaming inserts are at-least-once, not exactly-once. When Dataflow's insert call times out, the bundle is retried; if the original insert actually committed, the retry writes the same window's rows again, producing duplicates. Idempotent deduplication keys or BigQuery Storage Write API with exactly-once semantics prevent this.
- ✗
The pipeline uses default triggers instead of after-watermark triggers.
Why it's wrong here
Default triggers fire repeatedly on early and late panes, emitting multiple results per window; after-watermark triggers emit once when the watermark passes, preventing duplicates. Default triggering suits speculative low-latency output where downstream deduplication is acceptable, not exactly-once BigQuery writes.
- ✗
The fixed window duration is too short, causing overlapping windows.
Why it's wrong here
Fixed windows of equal length never overlap; each element belongs to exactly one window, so shortening the duration cannot create duplicates. Duplicates arise from repeated trigger firings per window. Short windows suit low-latency dashboards, but they change emission frequency, not row duplication.
- ✗
The pipeline is using too many Dataflow workers, causing load balancing issues.
Why it's wrong here
Worker count affects parallelism and throughput, not result correctness; Dataflow's shuffle and checkpointing keep per-key state consistent regardless of worker scaling. Duplicate rows stem from trigger behaviour emitting multiple panes. Adding workers suits backlog reduction during peak load, not fixing duplicate output.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.