Courseiva
Analyzing Queries →mediumMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

A data analyst notices that a query performing a join between a large fact table and a small dimension table is slow. The analyst wants to ensure that the small table is broadcasted to all executors to avoid a shuffle. Which configuration should the analyst adjust to increase the likelihood of a broadcast join?

⚠ Common exam trap

The trap here is thinking that enabling adaptive query execution alone guarantees a broadcast join, when it still depends on the autoBroadcastJoinThreshold setting.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

spark.sql.autoBroadcastJoinThreshold

The key to encouraging a broadcast join is to ensure the small table's size is below the autoBroadcastJoinThreshold. By increasing this threshold, the analyst allows Spark to consider the small table for broadcast, eliminating the shuffle of the large table and significantly speeding up the join.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    spark.sql.autoBroadcastJoinThreshold

    Why this is correct

    This configuration sets the maximum size (in bytes) of a table that can be broadcasted in a join. By increasing this threshold, the analyst allows larger tables to be considered for broadcast, which can eliminate the shuffle of the large table. In this scenario, if the small dimension table's size is below the threshold, it will be broadcasted, improving join performance.

  • ✗

    spark.sql.adaptive.enabled

    Why it's wrong here

    Enabling adaptive query execution allows Spark to optimize the query plan at runtime, including converting sort-merge joins to broadcast joins if a table is small enough. However, it does not directly increase the likelihood of broadcast; it relies on the autoBroadcastJoinThreshold. If that threshold is too low, AQE might not choose broadcast. So this alone is not sufficient.

  • ✗

    spark.sql.shuffle.partitions

    Why it's wrong here

    This setting controls the number of partitions used when shuffling data for joins or aggregations. While increasing it can improve parallelism, it does not trigger a broadcast join. In fact, a broadcast join avoids shuffle altogether. Adjusting this parameter might help other shuffle-heavy operations but does not address the goal of broadcasting the small table.

  • ✗

    spark.sql.broadcastTimeout

    Why it's wrong here

    This configuration sets the timeout for broadcasting a table. Increasing it can prevent timeouts when broadcasting large tables, but it does not influence the decision to broadcast. The decision is based on the table size relative to autoBroadcastJoinThreshold. Therefore, adjusting the timeout does not make a broadcast join more likely; it only affects the broadcast process once chosen.

About these practice questions

One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.