PDE Ingesting and Processing the Data Practice Question
Your team runs a Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. During a deployment, the pipeline is stopped and later restarted with the same job name but without draining. After the restart, downstream reports show duplicate rows for events that were processed just before the stop. You need the pipeline to resume without reprocessing already-published messages. What should you do?
⚠ Common exam trap
The trap here is assuming that restarting with the same job name or adding the update flag preserves exactly-once progress, when only draining commits in-flight Pub/Sub acknowledgements.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Drain the pipeline before stopping it, then start the replacement pipeline from the saved state.
Stopping a streaming job without draining leaves in-flight Pub/Sub messages unacknowledged, so they are redelivered when a replacement job starts, producing duplicates. Draining first lets the pipeline finish processing and commit those acknowledgements and BigQuery writes, so the replacement resumes from a consistent state and avoids reprocessing the same events.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of workers with --maxNumWorkers before restarting the pipeline.
Why it's wrong here
Worker count affects throughput and parallelism, not exactly-once semantics or checkpoint recovery. Adding workers before a restart does not change which Pub/Sub messages are considered acknowledged. The duplicates arise because uncommitted messages are replayed, so scaling the worker pool leaves the reprocessing behavior unchanged.
- ✓
Drain the pipeline before stopping it, then start the replacement pipeline from the saved state.
Why this is correct
Draining a streaming Dataflow job stops ingestion of new data, finishes processing the in-flight elements, and commits the Pub/Sub acknowledgements and BigQuery writes. When the replacement job starts, it resumes from the committed state rather than replaying messages that were already acknowledged, which prevents the duplicate rows observed in the reports.
- ✗
Enable --enableStreamingEngine on the restarted pipeline to preserve message state.
Why it's wrong here
Streaming Engine changes how the pipeline executes, moving work off the worker VMs, but it does not by itself persist or restore Pub/Sub acknowledgement state across a stop and restart. The duplicate rows come from message replay, and enabling Streaming Engine does not prevent that replay from occurring.
- ✗
Restart the pipeline using the --update option so the transform graph is replaced in place.
Why it's wrong here
Using --update replaces the pipeline graph while keeping the same job, but it does not change the streaming checkpoint behavior for the messages that were already acknowledged. Messages processed before the stop remain committed, and the restarted graph can still reprocess unacknowledged messages, so duplicates are not eliminated by this flag alone.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.