Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
Your Spark application is experiencing severe data skew while performing a join between a large fact table and a small dimension table. Which technique should you apply to optimize performance?
⚠ Common exam trap
Candidates often suggest salting or repartitioning as the first step, ignoring that broadcasting is the most efficient and direct way to handle skew when a small table is involved.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Force a broadcast join for the small dimension table.
Broadcast hash joins are the most effective solution for skew caused by a large table join when one side is small enough to fit in memory. By broadcasting the small table to all executor nodes, Spark avoids the expensive shuffle operation that causes data skew. This prevents specific partitions from becoming hotspots, which is critical for maintaining stable and performant ETL pipelines in production Databricks environments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of partitions using repartition() on the large table.
Why it's wrong here
Increasing partitions on the large table alone does not resolve the imbalance if the skew exists in the join keys. Repartitioning may spread data differently but does not eliminate the underlying concentration of rows associated with specific keys, thus failing to mitigate the skew in a join operation.
- ✗
Enable AQE and increase the spark.sql.shuffle.partitions configuration.
Why it's wrong here
While Adaptive Query Execution is powerful, simply increasing shuffle partitions does not resolve skew. AQE requires specific enabling flags for skew join optimization to handle unbalanced partition sizes effectively. Increasing shuffle partitions globally can actually hurt performance by creating too many small tasks for the executor overhead.
- ✓
Force a broadcast join for the small dimension table.
Why this is correct
Broadcasting the smaller DataFrame eliminates the need for a sort-merge join, which is where skewed data causes bottlenecks. By copying the small table to every executor, you perform a map-side join, effectively avoiding the shuffle of the large table and preventing skewed keys from overloading specific worker nodes.
- ✗
Apply a salting technique to the join keys of the small table.
Why it's wrong here
Salting is typically applied to the skewed keys of the large table, not the dimension table. Modifying the small dimension table's keys would lead to incorrect join results because the join condition would no longer match the original keys in the large table, resulting in data loss.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.