Courseiva
mediumMultiple Choice

PDE Practice Question: A data pipeline using Cloud Pub/Sub and Cloud…

A data pipeline using Cloud Pub/Sub and Cloud Dataflow is experiencing duplicate messages. The source system publishes messages at least once. What Dataflow technique ensures exactly-once processing?

⚠ Common exam trap

A common mix-up: candidates confuse 'exactly-once processing' with 'exactly-once delivery' from the source, but Pub/Sub only guarantees at-least-once delivery, so the responsibility for deduplication falls on the Dataflow pipeline and its sink design, not on windowing or engine settings.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use idempotent sinks

Idempotent sinks ensure that even if Cloud Pub/Sub delivers the same message multiple times (due to its at-least-once delivery semantics), the Dataflow pipeline can deduplicate or safely reapply the same data without causing duplicates in the output. This is achieved by designing the sink (e.g., BigQuery with insertId, Cloud Storage with unique filenames) to recognize and ignore repeated writes, effectively providing exactly-once processing semantics downstream.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use idempotent sinks

    Why this is correct

    Idempotent sinks deduplicate writes at the destination, so repeated delivery of the same Pub/Sub message produces one persisted record. This satisfies the exactly-once requirement despite at-least-once publication, because Dataflow's streaming engine pairs sink idempotency with per-message deduplication rather than relying on the source.

  • ✗

    Use GlobalWindows

    Why it's wrong here

    GlobalWindows places every element in one never-closing window, which aggregates across the whole stream rather than removing duplicates. It is tempting because a single window seems to consolidate records, but it changes grouping, not delivery semantics. Exactly-once instead relies on Dataflow deduplicating Pub/Sub messages by their unique message IDs.

  • ✗

    Set watermark threshold

    Why it's wrong here

    Watermark thresholds govern when windows fire for late data; they do not deduplicate records. It is tempting because watermarks shape streaming correctness, but they address event-time completeness, not repeated delivery. Exactly-once requires Dataflow's built-in deduplication keyed on Pub/Sub message IDs, which watermarks cannot provide.

  • ✗

    Enable streaming engine

    Why it's wrong here

    The streaming engine alters execution, autoscaling and latency characteristics of pipelines; it does not itself deduplicate at-least-once Pub/Sub deliveries. It is tempting because it is a streaming-specific Dataflow feature, yet enabling it leaves duplicate records intact. Exactly-once depends on Dataflow's ID-based deduplication of Pub/Sub messages.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.