Courseiva

Databricks-Spark-Assoc Spark Architecture and Components Practice Question

A Spark application running on Databricks uses broadcast joins for a small dimension table. A developer notices that the broadcast variable is not being sent to executors as expected, causing a shuffle instead. Which configuration property directly controls the maximum size of a table that Spark will automatically broadcast?

⚠ Common exam trap

The trap here is conflating properties that affect broadcast execution (block size, timeout) with the one that controls the size-based decision to broadcast, which is autoBroadcastJoinThreshold.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

spark.sql.autoBroadcastJoinThreshold

The automatic selection of a broadcast join is governed by spark.sql.autoBroadcastJoinThreshold, which defaults to 10 MB in Spark. If the small table's estimated size is below this threshold, Spark broadcasts it; otherwise, it uses a shuffle join. Other properties like shuffle partitions, broadcast block size, and broadcast timeout affect execution details but not the decision to broadcast based on table size.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    spark.sql.autoBroadcastJoinThreshold

    Why this is correct

    spark.sql.autoBroadcastJoinThreshold sets the maximum size in bytes of a table that Spark will broadcast automatically during a join. If the small table's size exceeds this threshold, Spark falls back to a shuffle join. Adjusting this property directly influences whether a broadcast join is chosen, making it the correct control for the scenario.

  • ✗

    spark.broadcast.blockSize

    Why it's wrong here

    spark.broadcast.blockSize controls the size of blocks used when transferring broadcast data, affecting how the broadcast is chunked, not whether a table is eligible for broadcasting. It does not set a size limit for automatic broadcast joins. Changing this property may influence transfer efficiency but will not prevent a shuffle join when the threshold is exceeded.

  • ✗

    spark.sql.broadcastTimeout

    Why it's wrong here

    spark.sql.broadcastTimeout specifies how long Spark waits for a broadcast to complete before timing out, in seconds. It does not determine the maximum size of a table eligible for broadcasting. While it can affect broadcast join reliability, it is not the property that controls whether a broadcast join is chosen based on table size.

  • ✗

    spark.sql.shuffle.partitions

    Why it's wrong here

    spark.sql.shuffle.partitions determines the number of partitions used when shuffling data for joins or aggregations, but it does not govern broadcast join selection. Even with a high value, Spark will not broadcast a table based on this setting; it only affects the parallelism of the shuffle. Therefore, it is unrelated to the broadcast size threshold.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.