Courseiva

PDE Ingesting and Processing the Data Practice Question

Your team runs a Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. During a deployment, the pipeline is stopped and later restarted with the same job name but without draining. After the restart, downstream reports show duplicate rows for events that were processed just before the stop. You need the pipeline to resume without reprocessing already-published messages. What should you do?

⚠ Common exam trap

The trap here is assuming that restarting with the same job name or adding the update flag preserves exactly-once progress, when only draining commits in-flight Pub/Sub acknowledgements.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Drain the pipeline before stopping it, then start the replacement pipeline from the saved state.

Stopping a streaming job without draining leaves in-flight Pub/Sub messages unacknowledged, so they are redelivered when a replacement job starts, producing duplicates. Draining first lets the pipeline finish processing and commit those acknowledgements and BigQuery writes, so the replacement resumes from a consistent state and avoids reprocessing the same events.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of workers with --maxNumWorkers before restarting the pipeline.

    Why it's wrong here

    Worker count affects throughput and parallelism, not exactly-once semantics or checkpoint recovery. Adding workers before a restart does not change which Pub/Sub messages are considered acknowledged. The duplicates arise because uncommitted messages are replayed, so scaling the worker pool leaves the reprocessing behavior unchanged.

  • ✓

    Drain the pipeline before stopping it, then start the replacement pipeline from the saved state.

    Why this is correct

    Draining a streaming Dataflow job stops ingestion of new data, finishes processing the in-flight elements, and commits the Pub/Sub acknowledgements and BigQuery writes. When the replacement job starts, it resumes from the committed state rather than replaying messages that were already acknowledged, which prevents the duplicate rows observed in the reports.

  • ✗

    Enable --enableStreamingEngine on the restarted pipeline to preserve message state.

    Why it's wrong here

    Streaming Engine changes how the pipeline executes, moving work off the worker VMs, but it does not by itself persist or restore Pub/Sub acknowledgement state across a stop and restart. The duplicate rows come from message replay, and enabling Streaming Engine does not prevent that replay from occurring.

  • ✗

    Restart the pipeline using the --update option so the transform graph is replaced in place.

    Why it's wrong here

    Using --update replaces the pipeline graph while keeping the same job, but it does not change the streaming checkpoint behavior for the messages that were already acknowledged. Messages processed before the stop remain committed, and the restarted graph can still reprocess unacknowledged messages, so duplicates are not eliminated by this flag alone.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.