Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

Your team is using a shared cluster for development. A user reports that their job is slow because the cluster memory is frequently filled by large data broadcasts. What configuration adjustment should you make to prevent this issue across all jobs on the cluster?

⚠ Common exam trap

Candidates often try to increase cluster memory or change instance types. While this might temporarily fix the symptom, it doesn't address the root cause of Spark attempting to broadcast tables that are too large.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Decrease the 'spark.sql.autoBroadcastJoinThreshold'.

Controlling the broadcast threshold prevents Spark from automatically attempting to broadcast large tables that exceed the available memory, which is the most common cause of OOMs on shared clusters. Setting this value correctly ensures that only appropriately small tables are broadcast, forcing the engine to use a join shuffle instead. This maintains stability for all users on the shared cluster, preventing one user's job from impacting others.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase 'spark.driver.memory'.

    Why it's wrong here

    Increasing driver memory does not fix the issue, as the broadcast happens on the executor side where the join computation takes place. If the broadcast data is too large, it will still lead to executor-level OOM errors, regardless of how much memory is allocated to the driver node.

  • ✓

    Decrease the 'spark.sql.autoBroadcastJoinThreshold'.

    Why this is correct

    Reducing this threshold prevents the optimizer from choosing a broadcast join for tables that are larger than the specified limit. By forcing the engine to use a sort-merge join instead, you ensure that the memory stays within safe limits, preventing OOM errors on shared clusters where resources are finite.

  • ✗

    Increase the number of shuffle partitions.

    Why it's wrong here

    Increasing shuffle partitions helps with data parallelism but does not address the issue of broadcast join memory consumption. A broadcast join happens independently of shuffle partitions, so changing this setting will not prevent the executor from running out of memory when trying to hold the large broadcasted table.

  • ✗

    Enable 'spark.sql.adaptive.enabled'.

    Why it's wrong here

    Adaptive Query Execution (AQE) is helpful for optimization, but it can actually decide to convert a shuffle join into a broadcast join if it determines the table is small enough. In a shared cluster, this might still lead to OOMs if the threshold is too high, so manual thresholding is safer.

About these practice questions

One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.