Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A developer is tuning a Spark job that performs a join between a large fact table and a medium-sized dimension table. The job suffers from data skew, with a few keys having a disproportionately large number of rows. The developer wants to mitigate the skew. Which two actions are most effective? (Choose two.)
⚠ Common exam trap
The trap here is assuming that increasing shuffle partitions or repartitioning on the join key will fix skew, when they can actually worsen it by concentrating hot keys.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Adaptive Query Execution (AQE) and set spark.sql.adaptive.skewJoin.enabled to true.
Data skew during a join occurs when a few keys have many more rows than others, causing some tasks to run much longer. Adaptive Query Execution with skew join optimization dynamically splits skewed partitions, while manual salting distributes hot keys across multiple tasks. Both approaches balance the workload and reduce the impact of skew.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable Adaptive Query Execution (AQE) and set spark.sql.adaptive.skewJoin.enabled to true.
Why this is correct
AQE with skew join optimization automatically detects skewed partitions during the shuffle and splits them into smaller sub-partitions. This balances the workload across tasks without manual intervention. It is a built-in feature that directly addresses skew by dynamically handling oversized partitions, making it an effective solution.
- ✗
Broadcast the medium-sized dimension table to all executors.
Why it's wrong here
Broadcasting a medium-sized table can avoid a shuffle join if the table fits in memory, but it does not solve skew on the fact table side. If the fact table has skewed keys, the broadcast join may still result in uneven task durations because the large partitions remain. Broadcasting helps with small tables, but here the dimension is medium and the skew is on the fact side.
- ✗
Use a repartition on the join key before the join to evenly distribute the data.
Why it's wrong here
Repartitioning on the join key uses a hash partitioner, which sends all rows with the same key to the same partition. This actually exacerbates skew because all rows for a hot key land in one partition. Repartitioning does not solve skew; it can make it worse by concentrating the hot key. Other techniques like salting or AQE are needed.
- ✓
Manually salt the skewed keys in the fact table by adding a random suffix and replicate the dimension table accordingly.
Why this is correct
Salting involves adding a random prefix or suffix to skewed keys in the large table and replicating the smaller table to match. This distributes the skewed keys across multiple partitions, balancing the load. It is a manual but effective technique when AQE is not sufficient or unavailable. It directly addresses the uneven distribution.
- ✗
Increase spark.sql.shuffle.partitions to a very high number.
Why it's wrong here
Increasing shuffle partitions creates more, smaller partitions overall, but if a single key has a massive number of rows, that key's data will still be concentrated in one or a few partitions. Skew is about uneven key distribution, not total partition count. Simply adding partitions does not redistribute the skewed key's rows; it may even create more empty partitions.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.