hardMultiple Choice
PDE Practice Question: A Dataflow streaming pipeline processes events…
A Dataflow streaming pipeline processes events from Pub/Sub and writes to BigQuery using a dynamically generated table destination based on the event type. The pipeline is experiencing high latency, and the worker CPU utilization is low. Which action is most likely to reduce latency?
⚠ Common exam trap
A common mix-up: candidates assume low CPU utilization means workers are underutilized and should be scaled down (Option B), when in fact low CPU with high latency indicates a bottleneck in shuffle or state management that is not compute-bound.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Dataflow Streaming Engine to improve throughput and reduce latency.
Dataflow Streaming Engine moves state and computation from worker VMs to the backend service, reducing per-worker overhead and enabling better resource utilization. This directly addresses the symptom of high latency with low CPU utilization, which indicates workers are bottlenecked on shuffle or state management rather than compute.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the batch size parameter in the BigQuery sink to write larger batches.
Why it's wrong here
Increasing the batch size may reduce the number of BigQuery streaming inserts but could increase memory usage and latency per batch. The low CPU utilization suggests the bottleneck is elsewhere, likely in shuffle or state management, so this change is unlikely to help.
- ✗
Reduce the number of workers to increase CPU utilization per worker.
Why it's wrong here
Reducing the number of workers might increase CPU utilization per worker, but if the bottleneck is due to shuffle or state management (not compute), fewer workers could worsen latency. The symptom of low CPU and high latency indicates a need for Streaming Engine, not worker scaling.
- ✓
Enable Dataflow Streaming Engine to improve throughput and reduce latency.
Why this is correct
Enabling Dataflow Streaming Engine moves state and computation from worker VMs to the backend service, reducing per-worker overhead and alleviating shuffle/state bottlenecks. This directly addresses the low CPU utilization and high latency, improving throughput.
- ✗
Increase the worker disk size to reduce I/O wait time.
Why it's wrong here
Increasing worker disk size helps with disk I/O bottlenecks, but the issue here is not disk I/O; low CPU utilization with high latency typically points to a bottleneck in shuffle or state management, not disk operations.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.