Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A Databricks job joins a 500 GB sales table with a 300 GB returns table on a customer_id key. A few customer_id values account for a large fraction of rows on both sides, and the job fails with executor OOM during the join. The developer wants to distribute the hot keys across more partitions without changing the query logic. Which technique should be used?
⚠ Common exam trap
The trap here is believing that increasing shuffle partitions or repartitioning by the join key will split a hot key, when hash partitioning always routes all rows for one key to a single partition.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Adaptive Query Execution with skew join handling so Spark splits skewed partitions at runtime.
AQE skew join handling identifies partitions that are disproportionately large after the shuffle and splits them into smaller sub-partitions, each handled by a separate task. This spreads the hot customer_id values across executors and resolves the OOM. Raising shuffle partitions or repartitioning by the join key keeps each hot key in one partition, and broadcasting a 300 GB table is infeasible.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Broadcast the 300 GB returns table to every executor to avoid shuffling the skewed keys.
Why it's wrong here
A 300 GB table is far too large to broadcast; the broadcast relation must fit in executor memory, and the default threshold is 10 MB. Attempting to broadcast it would cause driver or executor OOM before the join even starts. Broadcasting is appropriate only for small dimension tables, not for a 300 GB fact-like table.
- ✓
Enable Adaptive Query Execution with skew join handling so Spark splits skewed partitions at runtime.
Why this is correct
AQE skew join handling detects partitions that are much larger than the median after the shuffle and splits them into smaller sub-partitions, each processed as a separate task. This distributes the hot customer_id values across multiple tasks and relieves the executor OOM without changing the join logic. It is the built-in mechanism for exactly this scenario.
- ✗
Increase spark.sql.shuffle.partitions to 8000 to spread the skewed keys across more tasks.
Why it's wrong here
Raising shuffle partitions increases the number of output partitions, but a single hot key still hashes to exactly one partition. All rows for that key land in the same task, so the OOM on that task persists. More partitions can even increase overhead without relieving the skew, because hash partitioning cannot split one key across reducers.
- ✗
Repartition both DataFrames by customer_id with 8000 partitions before the join.
Why it's wrong here
Repartitioning by customer_id still uses hash partitioning, so every row for a hot key hashes to the same partition. The skewed task remains oversized and still OOMs. Repartitioning changes the number of partitions but cannot split a single key, so it does not address the root cause of the skew.
Visual reference
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.