Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A developer needs to join `orders` (large) with `customers` (small, a few thousand rows) on `customer_id`. They want to broadcast the small table and confirm the broadcast actually took effect. Which combination of actions is correct?

⚠ Common exam trap

The trap here is assuming that increasing the broadcast threshold or repartitioning the small table guarantees a broadcast join, when only the plan confirms it.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Call `orders.join(broadcast(customers), "customer_id")` and verify by inspecting the physical plan for `BroadcastHashJoin`

Broadcasting a small table avoids shuffling the large one, and the `broadcast()` hint is the explicit way to request it. Verification matters because hints can be ignored when the table exceeds the threshold or when the join type is unsupported. Inspecting the physical plan for `BroadcastHashJoin` is the standard confirmation, making the combination of hint plus plan inspection the correct answer.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase `spark.sql.autoBroadcastJoinThreshold` to a very large value and rely on the optimizer without checking the plan

    Why it's wrong here

    Raising the threshold may cause Spark to broadcast tables that are too large, leading to driver out-of-memory errors. More importantly, relying on the optimizer without inspecting the plan gives no confirmation that a broadcast join actually occurred. The requirement includes verification, which this option omits.

  • ✗

    Call `customers.repartition(1)` before the join so the small table is in one partition

    Why it's wrong here

    Repartitioning the small table to one partition does not trigger a broadcast join; it simply forces a shuffle of that table to a single partition. The join would still be a shuffle-based join unless the size threshold is met. This adds cost without guaranteeing the intended broadcast behavior.

  • ✓

    Call `orders.join(broadcast(customers), "customer_id")` and verify by inspecting the physical plan for `BroadcastHashJoin`

    Why this is correct

    Wrapping the small DataFrame in `broadcast()` hints Spark to use a broadcast join, and inspecting the physical plan via `explain()` confirms whether `BroadcastHashJoin` appears. If the plan shows it, the hint was honored; if not, a shuffle join is occurring. This is the correct way to force and then verify a broadcast join.

  • ✗

    Use `orders.join(customers.hint("shuffle_hash"), "customer_id")` and check the plan for `ShuffledHashJoin`

    Why it's wrong here

    `shuffle_hash` is the opposite of what the developer wants; it forces a shuffle-based hash join rather than a broadcast. Verifying `ShuffledHashJoin` in the plan would confirm the wrong strategy. This option directly contradicts the requirement to broadcast the small table.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.