Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A Spark job performing a join between a 10 GB table and a 5 MB lookup table is running slowly, and the physical plan shows a SortMergeJoin. You want to avoid the shuffle. What should you do?
⚠ Common exam trap
The trap here is thinking that increasing shuffle partitions or repartitioning will avoid the shuffle, when in fact they only change how the shuffle is performed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set spark.sql.autoBroadcastJoinThreshold to a value larger than 5 MB, such as 10 MB, and ensure the small table is broadcast.
Broadcasting the small table eliminates the need to shuffle the large table, converting the join to a BroadcastHashJoin. This is achieved by ensuring the small table's size is below the auto broadcast join threshold or by using an explicit broadcast hint. The other options either do not change the join strategy or introduce unnecessary shuffles.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Repartition both DataFrames on the join key before the join.
Why it's wrong here
Repartitioning both DataFrames on the join key will still require a full shuffle of both datasets, including the large 10 GB table. This does not avoid the shuffle; it simply ensures co-location. While it can help with data skew, it is not the right solution when one side is very small and could be broadcast. The goal is to eliminate the shuffle entirely, not to optimize it.
- ✗
Increase spark.sql.shuffle.partitions to 2000.
Why it's wrong here
Increasing shuffle partitions affects the number of partitions after a shuffle but does not change the join strategy. The physical plan will still show SortMergeJoin because the small table is not being broadcast. More partitions might reduce per-partition size but adds overhead and does not eliminate the shuffle. This setting is irrelevant to choosing a broadcast join.
- ✗
Use a cross join and filter afterward.
Why it's wrong here
A cross join produces the Cartesian product of both tables, which is extremely expensive and incorrect for a lookup join. It would generate a massive number of rows and likely fail or run for a very long time. This is not a valid optimization and would not avoid a shuffle; it would create an even larger one. The correct approach is to broadcast the small table.
- ✓
Set spark.sql.autoBroadcastJoinThreshold to a value larger than 5 MB, such as 10 MB, and ensure the small table is broadcast.
Why this is correct
The auto broadcast join threshold controls the maximum size of a table that can be broadcast to all executors. The default is 10 MB, but if the small table is slightly above the threshold or statistics are missing, Spark may choose SortMergeJoin. Explicitly setting the threshold higher or using broadcast() hint forces a BroadcastHashJoin, eliminating the shuffle of the large table. This is the correct approach to avoid the shuffle in this scenario.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.