Courseiva
Analyzing Queries →mediumMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

A data analyst runs a Databricks SQL query that joins a 50 GB fact_sales Delta table with a 200 MB dim_product Delta table. The query takes 15 minutes. The analyst runs EXPLAIN FORMATTED and sees a SortMergeJoin instead of a BroadcastHashJoin. The analyst has already confirmed that the small table is not being broadcast. Which action should the analyst take to improve performance?

⚠ Common exam trap

The trap here is assuming that broadcast joins are automatically used for any small table, when in fact the default threshold is only 10 MB and must be increased for larger small tables.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase spark.sql.autoBroadcastJoinThreshold to a value larger than 200 MB, such as 250 MB, so the small table can be broadcast.

Increasing the auto broadcast join threshold above the size of the small table enables Spark to broadcast it, converting the SortMergeJoin to a BroadcastHashJoin. This avoids shuffling the large fact table, which is the main cost. The default threshold is 10 MB, so a 200 MB table is not broadcast unless the threshold is raised.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Add a broadcast hint to the large fact table so it is replicated to all executors.

    Why it's wrong here

    Broadcasting the large 50 GB fact table would cause out-of-memory errors and network congestion, as it would replicate the entire large table to every executor. Broadcast hints should be applied to the small table, not the large one. This action would worsen the performance problem.

  • ✓

    Increase spark.sql.autoBroadcastJoinThreshold to a value larger than 200 MB, such as 250 MB, so the small table can be broadcast.

    Why this is correct

    The default auto broadcast join threshold is 10 MB, so a 200 MB table is not broadcast. Increasing the threshold above 200 MB allows Spark to broadcast the small table, converting the SortMergeJoin into a BroadcastHashJoin. This eliminates the shuffle of the large fact table and significantly improves performance.

  • ✗

    Set spark.sql.autoBroadcastJoinThreshold to -1 to disable broadcasting and force a shuffle hash join.

    Why it's wrong here

    Setting the auto broadcast join threshold to -1 disables broadcast joins entirely, which would prevent the small dimension table from being broadcast. This would not improve performance; it would likely make the join slower by forcing a shuffle. The analyst wants to enable broadcast for the small table, not disable it.

  • ✗

    Repartition both tables on the join key before the join to ensure co-located partitions.

    Why it's wrong here

    While repartitioning can help with data skew, it still requires a full shuffle of both tables. The small table is only 200 MB and can easily be broadcast, avoiding the shuffle of the large table entirely. Repartitioning would not address the root cause and would add unnecessary shuffle overhead.

About these practice questions

Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.