PDE Ingesting and Processing the Data Practice Question
A financial company uses a Dataflow streaming pipeline to read transactions from Pub/Sub and write to BigQuery. They need exactly-once processing semantics for the BigQuery writes and want to avoid duplicates during pipeline updates. Which approach should they use?
⚠ Common exam trap
The trap here is assuming that insertId deduplication with streaming inserts guarantees exactly-once, when it only provides best-effort deduplication.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the BigQuery Storage Write API with exactly-once semantics in the Dataflow BigQueryIO connector.
The Storage Write API with exactly-once semantics in BigQueryIO provides deduplication through stream offsets and supports pipeline updates without duplicates. It is the designed solution for exactly-once streaming writes to BigQuery in Dataflow, unlike legacy streaming inserts or file loads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the BigQuery Storage Write API with exactly-once semantics in the Dataflow BigQueryIO connector.
Why this is correct
The BigQueryIO connector can use the Storage Write API with exactly-once semantics, which deduplicates writes using stream offsets and supports seamless pipeline updates without duplicates. This is the recommended approach for exactly-once streaming writes to BigQuery in Dataflow.
- ✗
Use BigQueryIO with STREAMING_INSERTS and enable insertId-based deduplication.
Why it's wrong here
STREAMING_INSERTS with insertId provides best-effort deduplication, but it is not guaranteed exactly-once, especially across pipeline updates or retries. The legacy streaming API can also incur higher costs and does not offer the same exactly-once guarantees as the Storage Write API.
- ✗
Use BigQueryIO with the FILE_LOADS method and trigger frequent load jobs.
Why it's wrong here
FILE_LOADS writes data to Cloud Storage and then loads it into BigQuery, which is efficient for batch but not designed for low-latency streaming. It does not provide exactly-once semantics for streaming updates in the same way and introduces latency, so it does not meet the requirement.
- ✗
Write to a temporary table and run a MERGE statement periodically to deduplicate.
Why it's wrong here
A MERGE-based deduplication adds complexity and latency, and it requires an additional scheduling mechanism. It does not provide exactly-once semantics at the pipeline level and can still produce duplicates if the merge is not coordinated with the stream, so it is not the recommended approach.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.