mediumMultiple Select
PDE Practice Question: Which TWO actions are recommended to improve the…
Which TWO actions are recommended to improve the reliability of a Cloud Dataflow streaming pipeline that processes event data from Pub/Sub?
⚠ Common exam trap
Many exam-takers confuse reliability with throughput or latency, and may incorrectly choose micro-batching or disabling autoscaling as reliability improvements, when in fact Dataflow's reliability comes from its managed backend services like Streaming Engine.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Dataflow Streaming Engine.
Option B is correct because Dataflow Streaming Engine moves pipeline state and shuffle execution out of the worker VMs and into the Dataflow service, which improves reliability by decoupling state management from worker lifecycle events such as restarts, autoscaling, and updates. Option C is correct because enabling exactly-once processing sinks (for example, BigQuery with row-level deduplication via insertion IDs) prevents duplicate or lost records when retries occur, which is essential for reliable event processing from Pub/Sub. Option A is not recommended because a 10-second acknowledgment deadline is too short for many streaming workloads and would cause premature redelivery and duplicate processing rather than improving reliability. Option D is incorrect because disabling autoscaling does not improve reliability and can actually reduce throughput and resilience during load spikes. Option E is incorrect because micro-batching with a small batch size is not a Dataflow reliability best practice and can increase overhead and latency without addressing state or delivery guarantees.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a pull subscription with a 10-second acknowledgment deadline.
Why it's wrong here
A 10-second acknowledgment deadline is too short for Dataflow streaming pipelines, causing Pub/Sub to redeliver messages before processing completes and producing duplicates. It is tempting because pull subscriptions with tuned deadlines suit low-latency, high-throughput consumers, but streaming reliability needs a deadline exceeding worst-case processing time.
- ✓
Enable Dataflow Streaming Engine.
Why this is correct
Streaming Engine moves pipeline state and shuffle execution off the worker VMs to the Dataflow service, so workers no longer persist state to attached disks. This reduces worker restart and pipeline update disruption, improving reliability for Pub/Sub streaming pipelines.
- ✓
Enable exactly-once processing sinks (e.g., BigQuery with guaranteed row-level insertion).
Why this is correct
Exactly-once sinks prevent duplicate records when the pipeline retries bundles after worker failures or restarts. BigQuery's row-level insertion with deterministic deduplication ensures each Pub/Sub event is written once, preserving correctness under the at-least-once delivery that Pub/Sub guarantees.
- ✗
Disable autoscaling to prevent worker churn.
Why it's wrong here
Disabling autoscaling prevents Dataflow from adding workers during backlog spikes, increasing latency and risking pipeline failure under load. It is tempting because fixed worker pools avoid churn and suit predictable, steady-state workloads, but reliability for variable Pub/Sub event streams requires autoscaling to absorb bursts without dropping messages.
- ✗
Use micro-batch processing with a small batch size.
Why it's wrong here
Micro-batching adds latency and does not address Pub/Sub message retention or Dataflow checkpointing, so it cannot improve reliability against event loss. It is tempting because micro-batching reduces per-record overhead and suits throughput-sensitive batch-like workloads, but reliability here requires exactly-once processing and durable acknowledgements.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.