Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A developer needs to join `orders` (large) with `customers` (small, a few thousand rows) on `customer_id`. They want to broadcast the small table and confirm the broadcast actually took effect. Which combination of actions is correct?
⚠ Common exam trap
The trap here is assuming that increasing the broadcast threshold or repartitioning the small table guarantees a broadcast join, when only the plan confirms it.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call `orders.join(broadcast(customers), "customer_id")` and verify by inspecting the physical plan for `BroadcastHashJoin`
Broadcasting a small table avoids shuffling the large one, and the `broadcast()` hint is the explicit way to request it. Verification matters because hints can be ignored when the table exceeds the threshold or when the join type is unsupported. Inspecting the physical plan for `BroadcastHashJoin` is the standard confirmation, making the combination of hint plus plan inspection the correct answer.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase `spark.sql.autoBroadcastJoinThreshold` to a very large value and rely on the optimizer without checking the plan
Why it's wrong here
Raising the threshold may cause Spark to broadcast tables that are too large, leading to driver out-of-memory errors. More importantly, relying on the optimizer without inspecting the plan gives no confirmation that a broadcast join actually occurred. The requirement includes verification, which this option omits.
- ✗
Call `customers.repartition(1)` before the join so the small table is in one partition
Why it's wrong here
Repartitioning the small table to one partition does not trigger a broadcast join; it simply forces a shuffle of that table to a single partition. The join would still be a shuffle-based join unless the size threshold is met. This adds cost without guaranteeing the intended broadcast behavior.
- ✓
Call `orders.join(broadcast(customers), "customer_id")` and verify by inspecting the physical plan for `BroadcastHashJoin`
Why this is correct
Wrapping the small DataFrame in `broadcast()` hints Spark to use a broadcast join, and inspecting the physical plan via `explain()` confirms whether `BroadcastHashJoin` appears. If the plan shows it, the hint was honored; if not, a shuffle join is occurring. This is the correct way to force and then verify a broadcast join.
- ✗
Use `orders.join(customers.hint("shuffle_hash"), "customer_id")` and check the plan for `ShuffledHashJoin`
Why it's wrong here
`shuffle_hash` is the opposite of what the developer wants; it forces a shuffle-based hash join rather than a broadcast. Verifying `ShuffledHashJoin` in the plan would confirm the wrong strategy. This option directly contradicts the requirement to broadcast the small table.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.