Courseiva
Using Spark SQL →mediumMultiple Choice

Databricks-Spark-Assoc Using Spark SQL Practice Question

When performing a join between a very small table and a massive table, which Spark SQL optimization technique should be applied to prevent a full shuffle?

⚠ Common exam trap

Test-takers sometimes try to use partition pruning or caching hints instead of the specific broadcast hint required to eliminate shuffles during joins.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the /*+ BROADCAST(small_table) */ hint.

Broadcasting is a critical optimization technique for star-schema joins. By sending a copy of the small table to every executor, Spark avoids the expensive shuffle phase associated with shuffling the massive table. This significantly reduces network I/O and latency. For a Databricks developer, recognizing when to use hints or rely on the optimizer to perform broadcast joins is vital for writing performant, scalable SQL queries on large datasets.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the /*+ BROADCAST(small_table) */ hint.

    Why this is correct

    The BROADCAST hint explicitly instructs the Spark optimizer to broadcast the smaller table to all worker nodes. This eliminates the need for a shuffle, which is the most expensive part of a join, and is a standard way to ensure high-performance execution in Spark SQL for asymmetric join operations.

  • ✗

    Increase the shuffle partitions to 2000.

    Why it's wrong here

    Increasing the number of shuffle partitions does not prevent a shuffle from occurring. In fact, it increases the number of tasks, which may actually slow down a join if the data size is small. Proper optimization for joins involves reducing shuffling, not increasing the complexity of the shuffle process.

  • ✗

    Enable the 'autoBroadcastJoinThreshold' to a very low value.

    Why it's wrong here

    Setting the broadcast threshold to a low value will disable automatic broadcasting, potentially forcing Spark to perform a sort-merge join. This is the opposite of the desired outcome, as you want the small table to be broadcast to avoid the massive performance overhead of a standard shuffle join.

  • ✗

    Convert the massive table into an unmanaged view.

    Why it's wrong here

    Converting a table to a view has no impact on physical join execution strategies. A view is merely a logical pointer to an underlying query. The join execution plan remains determined by the Spark Catalyst optimizer based on table statistics, file size, and the query structure, not the view definition.

About these practice questions

This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.