hardMultiple Choice
Dataflow Keyed State OOM — Increase Workers Solution
A data pipeline ingests real-time events from Cloud Pub/Sub into BigQuery using Dataflow. The pipeline uses a sliding window of 5 minutes with a 1-minute period to aggregate event counts. Recently, the pipeline started failing with 'The worker failed to provide a heartbeat.' The Dataflow logs show high CPU usage on the workers. What is the best course of action to resolve the issue?
⚠ Common exam trap
Google Cloud often tests the misconception that reducing workers or changing window types is a universal fix for resource exhaustion, when in fact the immediate solution for heartbeat failures due to high CPU is to scale out the worker pool.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of workers and enable autoscaling to distribute the load.
The 'worker failed to provide a heartbeat' error combined with high CPU usage indicates that workers are overloaded and cannot process data fast enough to maintain their heartbeat to the Dataflow service. Increasing the number of workers and enabling autoscaling distributes the computational load across more machines, reducing per-worker CPU pressure and allowing heartbeats to be sent on time. This directly addresses the root cause of resource exhaustion.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Increase the number of workers and enable autoscaling to distribute the load.
Why this is correct
Scaling out workers with autoscaling directly relieves the CPU saturation causing missed heartbeats, since each worker's load shrinks as the sliding window's per-key aggregation is redistributed. This satisfies the stem's high-CPU constraint without altering window semantics, unlike tuning the 5-minute window or 1-minute period.
- ✗
Reduce the number of workers to minimize coordination overhead.
Why it's wrong here
Reducing workers lowers aggregate throughput, worsening the CPU saturation that causes missed heartbeats; scaling out adds capacity. It is tempting because fewer workers reduce shuffle and coordination traffic, which helps only when the bottleneck is inter-worker communication rather than per-worker CPU load.
- ✗
Use a global window with a trigger to reduce state size.
Why it's wrong here
A global window removes per-key windowing but accumulates unbounded state, worsening the CPU pressure behind the heartbeat failures. It is tempting because it reduces window bookkeeping, but a global window with triggers is correct when aggregating across all keys with no meaningful time boundaries.
- ✗
Change the windowing to a fixed 5-minute window to reduce computations.
Why it's wrong here
Fixed windows still aggregate per key and do not address the CPU saturation causing missed heartbeats; they merely alter emission timing. It is tempting because sliding windows compute overlapping panes, but fixed windows are correct when non-overlapping, aligned aggregates are genuinely required by the business logic.
Visual reference
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PDE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. Your company uses Cloud Dataflow to process streaming data from Pub/Sub. The pipeline occasionally fails with a 'worker terminated unexpectedly' error. What is the most likely cause of this error?
easy- ✓ A.Insufficient memory per worker causing OOM errors
- B.Incorrect VPC firewall rules blocking internal communication
- C.Staging location bucket lacks write permissions
- D.Pub/Sub subscription throughput quota exceeded
Why A: The 'worker terminated unexpectedly' error in Cloud Dataflow typically indicates that a worker process ran out of memory (OOM) and was killed by the operating system. This occurs when the pipeline's memory requirements exceed the configured worker machine type's memory capacity, often due to large windowing accumulations, skewed data, or inefficient state handling.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.