Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A job joining a large fact table with a small dimension table runs out of memory on executors during the join. The dimension table is about 40 MB after filtering and the configured spark.sql.autoBroadcastJoinThreshold is 10 MB. The join key is highly skewed in the fact table. Which action is most appropriate?
⚠ Common exam trap
The trap here is treating a large-to-small join as a skew problem to be salted, when broadcasting the small side removes the shuffle entirely.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Raise spark.sql.autoBroadcastJoinThreshold above the dimension table size so the small side is broadcast and no shuffle of the fact table occurs.
The dimension table is small enough to broadcast once the threshold is raised, which converts the join into a map-side operation and removes the shuffle of the large, skewed fact table. That eliminates the executor memory pressure caused by concentrating hot-key rows in a few reduce tasks. Salting and repartitioning address skew only in large-to-large joins and add cost here, while adding executor memory merely postpones the failure without changing the underlying plan.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Raise spark.sql.autoBroadcastJoinThreshold above the dimension table size so the small side is broadcast and no shuffle of the fact table occurs.
Why this is correct
Broadcasting the filtered 40 MB dimension table eliminates the shuffle of the large fact table and turns the join into a map-side operation. This avoids the memory pressure and shuffle associated with a sort-merge join on a skewed key. The threshold is a size guardrail, so raising it to cover the known small side is the intended control. This directly removes the source of the executor memory failure.
- ✗
Repartition the fact table by the join key with a high partition count before the join to distribute the skew.
Why it's wrong here
Repartitioning by the join key still sends all rows for a hot key to the same partition, so the skew remains concentrated. It adds a full shuffle of the large fact table while the small dimension is not broadcast. This increases network and disk cost without solving the memory spike. Broadcasting the small side is the more direct remedy.
- ✗
Increase spark.executor.memory so each executor can hold the skewed partitions of the fact table during the sort-merge join.
Why it's wrong here
Adding executor memory may delay the failure but does not remove the shuffle or the skew that causes a few partitions to be enormous. A single skewed key can still overwhelm one task regardless of total executor memory. This approach increases cost without fixing the root cause. The small dimension table makes a broadcast join the better structural fix.
- ✗
Salt the join key on both sides of the join so the skewed values are spread across many partitions.
Why it's wrong here
Salting is a valid technique for skewed joins between two large tables, but it requires expanding keys and a second aggregation, adding complexity and cost. Here the other side is only 40 MB, so broadcasting is simpler and cheaper. Salting would also need careful handling to avoid incorrect results. It is overkill for a large-to-small join.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.