Courseiva

Databricks-DA-Assoc Executing Queries with Databricks SQL Practice Question

A data analyst is troubleshooting a slow Databricks SQL query that joins a large fact table with a small dimension table. The query plan shows a broadcast hash join, but the analyst notices that the small table is not being broadcast as expected. Which configuration should the analyst check to ensure the small table is broadcast?

⚠ Common exam trap

The trap here is assuming that enabling adaptive query execution alone guarantees a broadcast join, when the auto-broadcast threshold must also be satisfied.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

spark.sql.autoBroadcastJoinThreshold

The broadcast hash join is governed by the autoBroadcastJoinThreshold configuration, which defines the maximum size for a table to be broadcast. If the small table exceeds this threshold, it will not be broadcast. Checking and potentially increasing this threshold can enable the broadcast, improving performance. Other settings affect different aspects of query execution and do not directly control broadcast eligibility.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    spark.sql.adaptive.enabled

    Why it's wrong here

    This setting enables adaptive query execution, which can dynamically optimize joins at runtime, including converting sort-merge joins to broadcast joins if one side is small enough. However, it relies on the auto-broadcast threshold. If the threshold is too low, adaptive execution may still not broadcast. The analyst should first check the threshold, making this option less direct.

  • ✓

    spark.sql.autoBroadcastJoinThreshold

    Why this is correct

    This Spark configuration sets the maximum size in bytes for a table to be considered for broadcasting in a join. If the small table's size exceeds this threshold, it will not be broadcast. The analyst should verify this value and adjust it if necessary to allow the small table to be broadcast, which can improve join performance by avoiding a shuffle.

  • ✗

    spark.databricks.delta.optimizeWrite.enabled

    Why it's wrong here

    This configuration controls whether Delta Lake optimizes write operations by reducing the number of small files. It has no impact on join strategies or broadcast behavior. The analyst's issue is about query execution, not write optimization, so this setting is irrelevant.

  • ✗

    spark.sql.shuffle.partitions

    Why it's wrong here

    This configuration controls the number of partitions used when shuffling data for joins or aggregations. While it affects performance, it does not determine whether a table is broadcast. Even with an optimal partition count, a table will not be broadcast if it exceeds the auto-broadcast threshold. Thus, it is not the correct setting to check for broadcast behavior.

About these practice questions

One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.