hardMultiple Choice
PDE Practice Question: A company's Dataflow pipeline uses the PubSubIO…
A company's Dataflow pipeline uses the PubSubIO source to read messages and writes to BigQuery via the BigQueryIO sink. The pipeline is running in Streaming mode with exactly-once semantics enabled. Occasionally, duplicate rows appear in BigQuery. What is the most likely reason?
⚠ Common exam trap
Google Cloud often tests the misconception that exactly-once semantics in Dataflow automatically deduplicates at the sink, but in reality, BigQuery requires explicit user-provided record IDs for deduplication during streaming inserts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The user-provided record ID for deduplication in BigQuery's streaming inserts is not being set for all messages, leading to duplicate rows.
In Dataflow streaming pipelines with exactly-once semantics, BigQuery's streaming inserts use user-provided record IDs for deduplication. If the record ID is not set for all messages, BigQuery cannot identify duplicates, and retries or redeliveries from Pub/Sub can result in duplicate rows. This is the most common cause of duplicates in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The user-provided record ID for deduplication in BigQuery's streaming inserts is not being set for all messages, leading to duplicate rows.
Why this is correct
BigQueryIO's streaming exactly-once deduplication relies on a deterministic record ID per message; when that ID is absent or non-unique, BigQuery cannot deduplicate retried inserts, so duplicates appear. The pipeline's exactly-once guarantee depends on this ID being set for every message.
- ✗
The pipeline is using the WriteResult method with WRITE_APPEND in batch mode, which can cause duplicates if retries happen.
Why it's wrong here
The stem specifies streaming mode with exactly-once semantics, so batch WRITE_APPEND retry behaviour is not the mechanism at play. It is tempting because WRITE_APPEND genuinely can duplicate rows under retries in batch pipelines, making it the right answer when exactly-once is disabled or the job runs in batch.
- ✗
The pipeline is experiencing the 'dataflow streaming log processing' bug, causing duplicate logs to be written.
Why it's wrong here
No such Dataflow bug exists; streaming log processing is a normal pipeline stage, not a duplication source. It is tempting because vague platform defects can appear to explain intermittent symptoms, but the actual cause lies in BigQueryIO's insertId-based deduplication window or Pub/Sub redelivery, not a fabricated logging fault.
- ✗
The PubSubIO source is configured with a dead-letter queue and messages are being redelivered without proper deduplication.
Why it's wrong here
Dead-letter queues receive messages that repeatedly fail processing; they do not cause redelivery of successfully processed messages into the main subscription. It is tempting because DLQs do involve redelivery mechanics, and would be relevant when diagnosing poison messages, but Pub/Sub at-least-once redelivery is the actual duplication source here.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.