Courseiva
Using Spark SQL →hardMultiple Choice

Databricks-Spark-Assoc Using Spark SQL Practice Question

A developer is using Spark SQL to join two DataFrames: a large fact table 'orders' and a smaller dimension table 'customers'. The join key is customer_id. The developer notices that the join is causing a shuffle and wants to avoid it. Which Spark SQL technique should be used to eliminate the shuffle for this join?

⚠ Common exam trap

Many candidates confuse the autoBroadcastJoinThreshold setting: setting it to -1 disables broadcast joins, while increasing it enables broadcasting for larger tables.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a broadcast hint: SELECT /*+ BROADCAST(customers) */ * FROM orders JOIN customers ON orders.customer_id = customers.customer_id

The broadcast hint instructs Spark to send the smaller table to all executors, so the large table can be joined locally without shuffling. This is the optimal approach when one side of the join is small enough to fit in memory. Other options either disable broadcast, use an invalid hint, or introduce additional shuffles. The broadcast join is a key optimization in Spark SQL for star-schema joins.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a broadcast hint: SELECT /*+ BROADCAST(customers) */ * FROM orders JOIN customers ON orders.customer_id = customers.customer_id

    Why this is correct

    The broadcast hint tells Spark to broadcast the smaller 'customers' table to all executors, avoiding a shuffle of the large 'orders' table. This is the correct technique when one side of the join is small enough to fit in memory. Spark SQL supports the BROADCAST hint, and it is the standard way to eliminate shuffle for such joins. The hint must be placed immediately after SELECT.

  • ✗

    Set spark.sql.autoBroadcastJoinThreshold to -1 to force broadcast joins.

    Why it's wrong here

    Setting autoBroadcastJoinThreshold to -1 disables broadcast joins entirely, which is the opposite of what is needed. The default threshold is 10MB; increasing it allows more tables to be broadcast. A value of -1 means no table will be broadcast, forcing a sort-merge join and shuffle. This would worsen performance for the small dimension table scenario.

  • ✗

    Repartition both DataFrames on customer_id before joining.

    Why it's wrong here

    Repartitioning both DataFrames on the join key would co-locate the data but still requires a full shuffle of both datasets to redistribute rows. This does not eliminate the shuffle; it adds an explicit shuffle step. While it can improve subsequent joins, it is not the technique to avoid shuffle for a small-large join. Broadcasting the small table is more efficient.

  • ✗

    Use a merge hint: SELECT /*+ MERGE(customers) */ * FROM orders JOIN customers ON orders.customer_id = customers.customer_id

    Why it's wrong here

    There is no MERGE hint in Spark SQL for joins. The MERGE hint is not a valid join strategy hint; Spark supports BROADCAST, MERGE, SHUFFLE_HASH, and SHUFFLE_REPLICATE_NL hints, but MERGE is used for merge joins (sort-merge) which still require a shuffle. Using an invalid hint may be ignored or cause an error. This does not eliminate the shuffle.

About these practice questions

This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.