Courseiva
hardMultiple Select

PDE Practice Question: Which THREE best practices should be followed…

Which THREE best practices should be followed when designing a Dataflow pipeline for real-time data processing?

⚠ Common exam trap

Google Cloud often tests the misconception that static side inputs are acceptable for streaming pipelines, but they are only appropriate for batch or bounded data; real-time pipelines require side inputs that can be periodically refreshed (e.g., via a streaming source or a periodic lookup).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set up monitoring alerts for system lag and data freshness.

Option A is correct because monitoring system lag (the difference between event time and processing time) and data freshness is essential for detecting pipeline backlog and ensuring timely real-time results in Dataflow. Option C is correct because watermark estimation is how Dataflow tracks event-time completeness and determines when to fire window results while gracefully handling late-arriving data. Option E is correct because idempotent sinks allow retries and replays to produce the same result, which is required to achieve exactly-once processing semantics in streaming pipelines. Option B is not appropriate because static side inputs loaded once at pipeline start cannot reflect changing reference data in a real-time stream, where dynamic or periodically refreshed side inputs are needed. Option D is not a best practice because global windows with early triggers discard windowing structure and can produce incorrect or incomplete aggregations for unbounded data, whereas proper event-time windowing with allowed lateness is preferred.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Set up monitoring alerts for system lag and data freshness.

    Why this is correct

    Real-time pipelines accumulate lag between event ingestion and output, so alerting on system lag and data freshness surfaces backpressure or stalled processing before downstream consumers act on stale results. This satisfies the stem's real-time processing requirement by detecting latency regressions continuously rather than after batch completion.

  • ✗

    Use static side inputs that are loaded once at pipeline start.

    Why it's wrong here

    Static side inputs are fixed at pipeline construction, so updated reference data never reaches a running streaming pipeline without redeployment. They suit bounded batch enrichment where lookup tables stay constant; real-time pipelines need periodically refreshed side inputs.

  • ✓

    Implement watermark estimation to handle late data.

    Why this is correct

    Event-time windows cannot know when all events have arrived, so watermark estimation defines when a window is complete and how long to wait for out-of-order records. This directly addresses late-arriving data, a defining constraint of real-time pipelines where network delays and mobile clients deliver events unpredictably.

  • ✗

    Use global windows with early triggers for low latency.

    Why it's wrong here

    Global windows never close without a trigger, so late or missing data produces incomplete, non-deterministic aggregates under continuous input. They suit bounded batch jobs; real-time processing needs windowing aligned to event time, such as fixed or sliding windows.

  • ✓

    Use idempotent sinks to ensure exactly-once processing.

    Why this is correct

    Dataflow delivers at-least-once by default, so retries after worker failures can duplicate writes. Idempotent sinks, such as upserts keyed on a deterministic identifier, make repeated writes harmless, achieving exactly-once semantics at the storage layer without relying on the runner alone.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.