Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
You are processing a large, highly skewed PySpark DataFrame in Databricks and want to optimize a forthcoming join operation against a small lookup dimension table. Which TWO strategies are valid and effective DataFrame API techniques to optimize this join performance? (Choose TWO)
⚠ Common exam trap
Candidates often assume broadcasting large fact tables or relying purely on automatic AQE join optimization solves all skew problems without manual intervention.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the broadcast() function on the small dimension table to distribute it to all worker nodes and avoid a shuffled hash join.
Optimizing joins involving skewed datasets requires avoiding shuffles for small tables or breaking up skewed keys. Broadcasting the small table eliminates the shuffle phase entirely, while salting keys distributes skewed join keys across multiple tasks, preventing memory bottlenecks on single executors during heavy shuffle exchanges.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the broadcast() function on the small dimension table to distribute it to all worker nodes and avoid a shuffled hash join.
Why this is correct
broadcast() ships the small dimension table to every executor, converting the join into a broadcast hash join and eliminating the shuffle of the large skewed DataFrame. This directly avoids the expensive shuffled hash join that skew would otherwise make prohibitively slow.
- ✗
Apply a broadcast join hint directly to the large skewed fact table to force executor nodes to cache its partitions in memory.
Why it's wrong here
Broadcasting a large fact table violates memory limits and triggers out-of-memory errors because driver nodes cannot gather and transmit massive datasets to executors. Broadcast hints must target small dimension tables that comfortably fit within executor memory.
- ✓
Introduce a salt column with random integers to the join keys of both DataFrames to evenly distribute skewed keys across multiple tasks.
Why this is correct
Salting splits heavy keys by appending a random integer, forcing Spark to distribute the workload across multiple tasks during the shuffle. The small table must be replicated accordingly to match the salted keys before executing the join.
- ✗
Increase the spark.sql.shuffle.partitions configuration to an extremely high number like 10000 to eliminate data skew entirely.
Why it's wrong here
Increasing shuffle partitions distributes tasks more finely but does not solve underlying data skew where a single key holds millions of rows. That specific key will still bottleneck whichever single task is assigned to process it.
- ✗
Convert the large DataFrame into a local Pandas DataFrame using toPandas() to perform the join operations locally on the driver node.
Why it's wrong here
toPandas() collects the entire dataset onto the driver's memory, causing out-of-memory failures and eliminating distributed parallelism, so it cannot optimise a large skewed join. It is tempting for small datasets where local Pandas joins are convenient, but broadcast joins or salting are the correct distributed techniques here.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.