PDE Ingesting and Processing the Data Practice Question
A media analytics company ingests clickstream events into Pub/Sub at a sustained rate of 2 GB/s. A Dataflow streaming pipeline reads these events, performs windowed aggregations, and writes results to BigQuery. The pipeline must handle occasional spikes up to 5 GB/s without data loss or excessive backlog. The operations team wants to minimize manual intervention and cost. What should you do to configure the Dataflow pipeline for dynamic scaling?
⚠ Common exam trap
The trap here is assuming that setting a high fixed number of workers is sufficient for spikes, but that ignores cost optimization and the benefits of dynamic scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling and set the maximum number of workers to 200, and use Streaming Engine to offload windowing and state management.
Autoscaling with a sufficient maximum worker count allows Dataflow to dynamically adjust to load, while Streaming Engine optimizes state management and reduces worker burden. This combination handles both sustained high throughput and spikes without manual tuning, and it minimizes cost during low periods. Fixed worker counts or batch processing would either be inefficient or fail to meet latency and scalability needs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the number of workers to a fixed value of 100 to handle peak load, and enable autoscaling with a maximum of 100 workers.
Why it's wrong here
Using a fixed worker count of 100 eliminates autoscaling benefits. During low traffic periods, you pay for idle workers, increasing cost. During spikes, if 100 workers are insufficient, the pipeline may lag. Dataflow's autoscaling is designed to adjust worker count dynamically based on backlog and CPU utilization, so fixing the count defeats the purpose and does not meet the requirement to minimize manual intervention and cost.
- ✓
Enable autoscaling and set the maximum number of workers to 200, and use Streaming Engine to offload windowing and state management.
Why this is correct
Enabling autoscaling allows Dataflow to add workers when backlog increases and remove them when load decreases, optimizing cost. Setting a high maximum ensures capacity for spikes. Streaming Engine moves pipeline state and windowing out of worker memory, improving scalability and reducing worker resource needs. This combination handles dynamic load without manual intervention and is the recommended approach for variable streaming workloads.
- ✗
Configure the pipeline to use a single worker with a high-memory machine type to reduce coordination overhead.
Why it's wrong here
A single worker cannot handle 2-5 GB/s throughput; it would become a bottleneck and cause severe lag or data loss. Vertical scaling with a high-memory machine does not provide the horizontal scalability needed for streaming spikes. Dataflow's strength is distributed processing, so using one worker is inappropriate and would not meet performance or reliability requirements.
- ✗
Use a batch pipeline with a trigger that runs every 5 minutes to process accumulated Pub/Sub messages.
Why it's wrong here
A batch pipeline introduces latency and does not provide continuous processing. Pub/Sub messages would accumulate, and the 5-minute interval may not handle spikes, leading to backlog. Batch pipelines also lack native streaming autoscaling based on real-time backlog. This approach fails to meet the requirement for handling sustained and spiky streaming data without data loss and with minimal manual intervention.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.