Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

You are writing a PySpark script to join two large tables. You want to ensure the join operation is optimized for performance by broadcasting the smaller table. Which configuration property should you adjust, or code construct should you use, to force this behavior?

⚠ Common exam trap

Candidates often try to manually set 'spark.sql.autoBroadcastJoinThreshold' to a massive value, which can cause driver OOM errors, instead of using the explicit 'broadcast()' hint on the specific dataframe.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.join(broadcast(small_df), 'id')

Broadcasting small tables is a critical optimization technique in Databricks to prevent expensive shuffles across the cluster. By utilizing the broadcast hint, developers explicitly instruct the Spark Catalyst optimizer to send the smaller table to all worker nodes. This minimizes data movement and significantly reduces latency during join operations. Understanding this mechanism is essential for building scalable ETL pipelines and ensuring efficient cluster resource utilization within the Databricks unified analytics platform.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    spark.conf.set('spark.sql.shuffle.partitions', '1')

    Why it's wrong here

    Adjusting shuffle partitions controls the degree of parallelism for wide transformations but does not force a broadcast join. Reducing this value to one can lead to memory overflow errors and severe performance degradation, as it forces all data onto a single executor rather than utilizing the cluster's distributed memory.

  • ✗

    spark.sql('SET spark.sql.autoBroadcastJoinThreshold = -1')

    Why it's wrong here

    This command disables the automatic broadcast join mechanism entirely. Setting the threshold to negative one instructs the optimizer to ignore automatic broadcasting, which is the opposite of the intended goal. This would lead to full sort-merge joins, increasing data shuffle operations across the network and reducing overall job efficiency.

  • ✓

    df.join(broadcast(small_df), 'id')

    Why this is correct

    The broadcast function provides a hint to the Catalyst optimizer to broadcast the specific dataframe to all worker nodes. This is the most efficient way to ensure a broadcast join occurs regardless of the default autoBroadcastJoinThreshold configuration, effectively eliminating the need for network-heavy shuffles during the join operation.

  • ✗

    df.repartition(100)

    Why it's wrong here

    Repartitioning changes the layout of the data across partitions based on the hash of the columns. While this can help balance data distribution, it triggers a full shuffle and does not influence the join strategy. It increases network overhead rather than optimizing the join execution plan through broadcasting.

About these practice questions

Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.