Databricks-DE-Pro Developing Code (Python/SQL) Practice Question
When running a PySpark job, you receive an 'Out of Memory (OOM)' error during a shuffle operation. Which configuration is the most appropriate to address this first?
⚠ Common exam trap
Candidates often choose increasing executor memory or cluster size first. While these might work, they are inefficient and costly compared to tuning shuffle partitions to reduce the memory footprint per task.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase spark.sql.shuffle.partitions
OOM errors during shuffles are frequently caused by partitions that are too large for the executor's memory. Increasing 'spark.sql.shuffle.partitions' is the primary corrective action, as it splits the data into a larger number of smaller partitions. This reduces the memory footprint of each individual task, allowing them to fit within the executor's heap memory and preventing the job from crashing due to memory exhaustion during data movement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase spark.driver.memory
Why it's wrong here
The driver memory handles metadata and task scheduling, not the actual data processing for shuffles. Increasing driver memory will not solve an OOM error that occurs during the execution of a shuffle on the worker nodes, as those tasks run in the executor memory space.
- ✓
Increase spark.sql.shuffle.partitions
Why this is correct
Increasing the number of shuffle partitions is the standard way to reduce the amount of data processed per task. By splitting the work across more partitions, each task consumes less memory, which helps resolve OOM issues occurring during shuffle-heavy operations like joins and aggregations.
- ✗
Decrease spark.executor.memory
Why it's wrong here
Decreasing executor memory will make OOM errors more likely, not less. Each task needs sufficient heap space to perform data processing, and reducing this allocation will force tasks to fail even earlier when handling large datasets during shuffle stages.
- ✗
Set spark.sql.shuffle.partitions to 1
Why it's wrong here
Setting shuffle partitions to 1 forces all data into a single partition, which would cause an OOM error for almost any dataset of non-trivial size. This is the worst possible configuration for managing memory in a distributed cluster and would lead to massive performance degradation.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.