PDE Ingesting and Processing the Data Practice Question
You are designing a Dataflow pipeline that needs to exactly-once process events from Pub/Sub and write to BigQuery using the Storage Write API. The pipeline may restart and could reprocess some messages. What setting ensures exactly-once semantics for the output?
⚠ Common exam trap
The trap is choosing legacy streaming inserts with insertId, which only provides best-effort deduplication, instead of the Storage Write API committed mode with Dataflow exactly-once, which is the actual mechanism for exactly-once semantics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the Storage Write API in committed mode and enable exactly-once semantic in Dataflow
The BigQuery Storage Write API in committed mode, combined with Dataflow's exactly-once processing, provides exactly-once semantics for writes to BigQuery. Committed mode uses a stream-based protocol where records are written and committed atomically, and Dataflow's exactly-once mode ensures that on pipeline restart, records are not duplicated. This is the recommended configuration for exactly-once streaming into BigQuery.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use the legacy streaming inserts with insertId for deduplication
Why it's wrong here
Legacy streaming inserts with insertId give best-effort deduplication within a short window and cannot guarantee exactly-once across pipeline restarts. It is tempting because insertId deduplication is the classic BigQuery streaming pattern, but the Storage Write API's exactly-once application-created stream is what the scenario requires.
- ✗
Use at-least-once delivery on Pub/Sub and idempotent writes to BigQuery
Why it's wrong here
At-least-once delivery plus idempotent writes yields effectively-once only if every write is truly idempotent; the Storage Write API's exactly-once mode instead uses application-created streams with offsets committed transactionally. It is tempting because idempotency is a common deduplication strategy, but it does not satisfy the stated exactly-once requirement.
- ✗
Use the Storage Write API in buffered mode with deduplication logic
Why it's wrong here
Buffered mode with manual deduplication still relies on application logic and cannot guarantee exactly-once across restarts; the Storage Write API's exactly-once mode uses application-created streams with committed offsets. It is tempting because buffered mode offers higher throughput, but that is a performance trade-off, not a correctness guarantee.
- ✓
Use the Storage Write API in committed mode and enable exactly-once semantic in Dataflow
Why this is correct
Committed mode on the Storage Write API makes each write atomic and idempotent via stream offsets, and enabling exactly-once in Dataflow deduplicates retried bundles. Together they prevent duplicate rows when the pipeline restarts and reprocesses Pub/Sub messages.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.