PDE Ingesting and Processing the Data Practice Question
You are using Dataproc to run a Spark job that reads data from Cloud Storage, performs aggregations, and writes results back to Cloud Storage. The job is failing with out-of-memory errors on the shuffle. Which optimization should you apply?
⚠ Common exam trap
PDE often tests the reflex to 'add more memory' for any OOM — candidates pick C, but shuffle OOM is a partitioning/skew problem, and the correct first lever is spark.sql.shuffle.partitions, not executor memory.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase spark.sql.shuffle.partitions
Out-of-memory errors during a shuffle in Spark are usually caused by too few shuffle partitions, which makes each partition (and its in-memory sort/aggregation buffer) too large. Increasing spark.sql.shuffle.partitions splits the shuffle into more, smaller partitions, reducing per-task memory pressure and allowing the job to complete.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Increase spark.sql.shuffle.partitions
Why this is correct
Raising `spark.sql.shuffle.partitions` splits shuffle data into more, smaller partitions, so each executor task holds less in memory during aggregation. This directly relieves the out-of-memory condition on the shuffle stage, satisfying the stem's constraint without changing cluster size or storage layout.
- ✗
Use RDDs instead of DataFrames
Why it's wrong here
RDDs discard the Catalyst optimiser and Tungsten code generation, so aggregations run as hand-written shuffle logic with higher serialisation overhead and no partition pruning. It is tempting when developers want fine-grained control over partitioning, which suits custom iterative algorithms rather than standard SQL-style aggregations.
- ✗
Increase spark.executor.memory
Why it's wrong here
Raising executor heap does not address shuffle spill, which is governed by the number of reduce partitions and per-partition record volume; too few partitions overload each task. It is tempting because executor memory tuning genuinely helps when executors themselves hold large cached datasets or wide transformations.
- ✗
Decrease the number of executors
Why it's wrong here
Fewer executors shrink total shuffle memory and parallelism, worsening the out-of-memory condition rather than relieving it. It tempts as a cost-saving tweak, but the fix is increasing executor memory or shuffle partitions so each task's shuffle data fits.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.