Courseiva

PDE Ingesting and Processing the Data Practice Question

You are building a Dataflow pipeline that reads from Pub/Sub, applies transformations, and writes to BigQuery. The pipeline must handle late-arriving data and ensure that the windowing and triggering are correct. Which THREE configurations should you consider? (Choose 3)

⚠ Common exam trap

A common misconception is that Dataflow Streaming Engine alone provides exactly-once processing. In reality, it is the combination of source/sink semantics (like the BigQuery Storage Write API) that ensures exactly-once, not the engine itself.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the BigQuery Storage Write API with committed mode to ensure exactly-once writes.

Option C is correct because the BigQuery Storage Write API in committed mode (exactly-once semantics) uses write streams with offsets and client-side deduplication, which is the recommended way to guarantee exactly-once writes from a Dataflow streaming pipeline into BigQuery. Option D is correct because setting an allowed lateness duration (via withAllowedLateness) on the windowing strategy tells Dataflow how long to retain window state and keep accepting late-arriving elements after the watermark passes the window end, which is exactly the requirement for handling late data. Option E is correct because configuring a triggering frequency (for example, AfterProcessingTime.pastFirstElementInPane().plusDelayOf(...) or a repeated trigger) controls how often intermediate or final results are emitted from each window, which directly governs the windowing/triggering behavior described in the scenario. Option A does not belong because Dataflow Streaming Engine is a runner feature that offloads state and windowing execution to the service for scalability and cost benefits; it does not itself provide exactly-once processing, which comes from the pipeline's sources, sinks, and deduplication logic. Option B does not belong because side inputs are used to enrich streaming records with additional (often static or slowly changing) data, which is unrelated to handling late-arriving data or to windowing and triggering correctness.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable Dataflow Streaming Engine for exactly-once processing.

    Why it's wrong here

    Streaming Engine offloads pipeline execution to the Dataflow service and improves autoscaling and latency, but exactly-once correctness comes from the runner's checkpointing and Pub/Sub acknowledgement model, not from enabling it. It is tempting because the feature is often recommended for streaming pipelines, yet it is not a windowing or triggering configuration.

  • ✗

    Use side inputs to enrich streaming data with static data.

    Why it's wrong here

    Side inputs join a streaming pipeline against a slowly changing reference dataset; they do not address windowing, triggers, or allowed lateness. The option is tempting because enrichment is a common streaming need, but the question asks specifically about handling late data and correct windowing semantics.

  • ✓

    Use the BigQuery Storage Write API with committed mode to ensure exactly-once writes.

    Why this is correct

    The BigQuery Storage Write API's committed mode provides exactly-once semantics through write-stream offsets, preventing duplicate rows when retries occur after late data triggers additional pane firings. This satisfies the pipeline's requirement for correct windowing and triggering, since speculative or repeated emissions from allowed lateness would otherwise insert duplicates into BigQuery.

  • ✓

    Set an allowed lateness duration to handle late-arriving data.

    Why this is correct

    Setting an allowed lateness duration lets the pipeline retain window state past the watermark, so late-arriving records are still assigned to their correct window and trigger a refined pane rather than being dropped. This directly satisfies the stem's requirement to handle late data while keeping windowing and triggering correct.

  • ✓

    Configure a triggering frequency to control how often results are emitted.

    Why this is correct

    Configuring a triggering frequency controls when pane results are emitted from a window, satisfying the requirement to handle late-arriving data. After-watermark triggers emit once the watermark passes the window end, while repeated early or late triggers emit speculative panes, letting downstream BigQuery writes reflect updates as further late elements arrive.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.