hardMultiple Choice
PDE Practice Question: A company runs a batch data processing workload…
A company runs a batch data processing workload using Dataproc clusters that are auto-scaled based on YARN memory utilization. During peak times, jobs take much longer than expected. Analysis shows the cluster is not scaling up despite high YARN memory utilization. What is the most likely cause?
⚠ Common exam trap
Many exam-takers assume autoscaling applies to all worker nodes equally, overlooking the Dataproc-specific distinction between primary and secondary workers and the autoscaler's limitation to secondary workers only.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The cluster is using primary workers only; auto-scaling only adds secondary workers
Dataproc clusters have two types of workers: primary workers (which run both HDFS and compute) and secondary workers (compute-only). The autoscaler can only add or remove secondary workers; it cannot scale primary workers. If the cluster uses only primary workers, the autoscaler has no secondary workers to add, so it cannot scale up even under high YARN memory utilization. This explains why the cluster remains static during peak times.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Spark dynamic allocation is disabled, preventing executors from using added workers
Why it's wrong here
Disabling Spark dynamic allocation leaves executor count fixed, but the autoscaler itself would still add YARN nodes; the reported symptom is that no workers are added at all. It is tempting because dynamic allocation genuinely governs executor elasticity, yet it cannot explain an autoscaler that never scales up.
- ✗
The cluster autoscaler is misconfigured to scale based on CPU, not memory
Why it's wrong here
Dataproc autoscaling reads YARN memory and pending containers, not CPU; the scaler cannot be switched to a CPU metric, so this misconfiguration is impossible. It tempts because CPU-based scaling is standard for stateless compute, and would fit a cluster sized by processor load rather than YARN memory pressure.
- ✗
The autoscaler is set to scale down secondary workers, not up
Why it's wrong here
Scale-down policy for secondary workers affects only removal of preemptible nodes and never governs scale-up decisions. It is tempting because secondary workers are the cheapest capacity and often the first place to look, but the autoscaler's upscaling is driven by YARN memory pressure on primary workers.
- ✓
The cluster is using primary workers only; auto-scaling only adds secondary workers
Why this is correct
Dataproc autoscaling adds only secondary workers to a cluster; primary worker count stays fixed. If the cluster runs primary workers alone, no scale-up occurs despite high YARN memory utilisation, which explains the stalled scaling during peak periods.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.