Courseiva

PDE Designing Data Processing Systems Practice Question

A retail company is designing a Dataflow pipeline to process point-of-sale transactions from Cloud Pub/Sub and write to BigQuery. The pipeline must handle late-arriving data up to 24 hours and ensure that all data is written to BigQuery exactly once, even in the event of worker failures. Which two features should the engineer implement to meet these requirements? (Choose two.)

⚠ Common exam trap

The trap here is assuming that streaming inserts with insertId provide exactly-once semantics, but they only offer best-effort deduplication and can still produce duplicates.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set the allowed lateness to 24 hours on the windowing strategy.

To handle late-arriving data up to 24 hours, the pipeline must set allowed lateness to 24 hours on the windowing strategy. To achieve exactly-once writes to BigQuery, the engineer should use the Storage Write API with a deterministic deduplication key and enable Dataflow's ExactlyOnce mode. Together, these features satisfy both requirements.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the Dataflow ExactlyOnce processing mode and enable the BigQueryIO.Write with FILE_LOADS and a trigger that fires after 24 hours.

    Why it's wrong here

    ExactlyOnce mode ensures no duplicates within Dataflow, but FILE_LOADS with a 24-hour trigger would delay writes and may not handle late data correctly. FILE_LOADS is for batch loads, not streaming, and a 24-hour trigger would cause high latency. This combination does not provide the required exactly-once semantics for streaming data with late arrivals.

  • ✓

    Set the allowed lateness to 24 hours on the windowing strategy.

    Why this is correct

    Setting allowed lateness to 24 hours allows the pipeline to accept and process data that arrives up to 24 hours after the window ends. This directly addresses the late-arriving data requirement. Combined with a proper trigger, it ensures that late data is included in the results and written to BigQuery.

  • ✓

    Enable the Dataflow ExactlyOnce processing mode and use the BigQueryIO.Write with STORAGE_WRITE_API and a deterministic deduplication key.

    Why this is correct

    The Storage Write API in BigQuery supports exactly-once semantics when used with a deterministic deduplication key. Combined with Dataflow's ExactlyOnce mode, this ensures that each record is written exactly once, even with retries. This meets the exactly-once requirement for writing to BigQuery.

  • ✗

    Use BigQueryIO.Write with the UseBeamSchema option and set the write disposition to WRITE_APPEND.

    Why it's wrong here

    UseBeamSchema and WRITE_APPEND simply append data to BigQuery. This does not provide exactly-once semantics; duplicates can occur if the pipeline retries. It also does not address late-arriving data or ensure that all data is written exactly once. This is a basic write configuration, not a solution for exactly-once.

  • ✗

    Configure the pipeline to use the BigQueryIO.Write with STREAMING_INSERTS and set the insertId based on a unique transaction identifier.

    Why it's wrong here

    STREAMING_INSERTS with insertId can provide best-effort deduplication, but it does not guarantee exactly-once semantics. BigQuery may still have duplicates if the insertId is not unique or if retries occur. Also, streaming inserts are not transactional and may not handle all failure scenarios. This does not meet the exactly-once requirement.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.