easyMultiple Choice
PDE Practice Question: Your team uses Cloud Dataproc to run a Spark ML…
Your team uses Cloud Dataproc to run a Spark ML training job. The job is failing with an error: 'Container killed by YARN for exceeding memory limits.' What should you do to fix this?
⚠ Common exam trap
Many candidates confuse scaling horizontally (adding nodes) with scaling vertically (increasing per-node resources), and assume more nodes will fix memory limits when the issue is per-container allocation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the spark.executor.memory property
The error 'Container killed by YARN for exceeding memory limits' indicates that the Spark executor process is using more memory than the YARN container allows. Increasing `spark.executor.memory` allocates a larger YARN container for each executor, providing the necessary headroom for the Spark application's memory demands, including overhead for off-heap memory and JVM internals.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Increase the spark.executor.memory property
Why this is correct
Raising `spark.executor.memory` gives each YARN container a larger heap, so the executor stays within the memory ceiling YARN enforces. The failure stems from the container exceeding its allocated limit, and this property directly governs that allocation, letting the Spark ML training job complete without being killed.
- ✗
Use preemptible VMs for faster execution
Why it's wrong here
Preemptible VMs change cost and availability, not the memory available to each YARN container, so the kill persists. It is tempting because preemptible workers genuinely reduce spend on fault-tolerant batch workloads, which would be the right choice if the goal were lowering cluster cost rather than resolving container memory limits.
- ✗
Increase the number of worker nodes
Why it's wrong here
Adding worker nodes increases total cluster memory but leaves the per-container allocation unchanged, so YARN still kills the container. It is tempting because scaling workers genuinely raises aggregate capacity for parallel tasks, which would be correct if the job needed more concurrent executors rather than a larger memory allocation per executor.
- ✗
Enable the external shuffle service
Why it's wrong here
The external shuffle service relocates shuffle data off executors; it does not raise the memory ceiling that YARN enforces per container. It is tempting because the shuffle service genuinely helps when executors are lost and shuffle data must persist, which would be correct for improving resilience on preemptible clusters rather than fixing memory limits.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.