Courseiva

PDE Ingesting and Processing the Data Practice Question

You are designing a Dataflow pipeline that needs to exactly-once process events from Pub/Sub and write to BigQuery using the Storage Write API. The pipeline may restart and could reprocess some messages. What setting ensures exactly-once semantics for the output?

⚠ Common exam trap

The trap is choosing legacy streaming inserts with insertId, which only provides best-effort deduplication, instead of the Storage Write API committed mode with Dataflow exactly-once, which is the actual mechanism for exactly-once semantics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Storage Write API in committed mode and enable exactly-once semantic in Dataflow

The BigQuery Storage Write API in committed mode, combined with Dataflow's exactly-once processing, provides exactly-once semantics for writes to BigQuery. Committed mode uses a stream-based protocol where records are written and committed atomically, and Dataflow's exactly-once mode ensures that on pipeline restart, records are not duplicated. This is the recommended configuration for exactly-once streaming into BigQuery.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the legacy streaming inserts with insertId for deduplication

    Why it's wrong here

    Legacy streaming inserts with insertId give best-effort deduplication within a short window and cannot guarantee exactly-once across pipeline restarts. It is tempting because insertId deduplication is the classic BigQuery streaming pattern, but the Storage Write API's exactly-once application-created stream is what the scenario requires.

  • ✗

    Use at-least-once delivery on Pub/Sub and idempotent writes to BigQuery

    Why it's wrong here

    At-least-once delivery plus idempotent writes yields effectively-once only if every write is truly idempotent; the Storage Write API's exactly-once mode instead uses application-created streams with offsets committed transactionally. It is tempting because idempotency is a common deduplication strategy, but it does not satisfy the stated exactly-once requirement.

  • ✗

    Use the Storage Write API in buffered mode with deduplication logic

    Why it's wrong here

    Buffered mode with manual deduplication still relies on application logic and cannot guarantee exactly-once across restarts; the Storage Write API's exactly-once mode uses application-created streams with committed offsets. It is tempting because buffered mode offers higher throughput, but that is a performance trade-off, not a correctness guarantee.

  • ✓

    Use the Storage Write API in committed mode and enable exactly-once semantic in Dataflow

    Why this is correct

    Committed mode on the Storage Write API makes each write atomic and idempotent via stream offsets, and enabling exactly-once in Dataflow deduplicates retried bundles. Together they prevent duplicate rows when the pipeline restarts and reprocesses Pub/Sub messages.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.