Databricks-DE-Pro Developing Code (Python/SQL) Practice Question
You are writing a Spark application that uses a Broadcast Hash Join. You want to force the join to use broadcast optimization for a specific table. How do you implement this in PySpark?
⚠ Common exam trap
Test-takers incorrectly try to use SQL hints or DataFrame configuration properties instead of the required PySpark functional wrapper to force a broadcast join.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
df1.join(broadcast(df2), 'key')
Using the 'broadcast()' function in PySpark explicitly hints to the Spark optimizer that a specific DataFrame should be broadcasted to all executors. This is highly effective when joining a large fact table with a small dimension table, as it avoids the expensive shuffle operation. Understanding this is crucial for performance tuning, as relying solely on the automatic broadcast threshold may lead to suboptimal join choices in complex queries.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
df1.join(broadcast(df2), 'key')
Why this is correct
The broadcast function wrapper forces the Spark optimizer to broadcast the smaller DataFrame (df2) to all nodes in the cluster. This avoids shuffling the larger table, which is a major performance boost for joins where one side is small enough to fit in the executor memory.
- ✗
df1.join(df2, 'key', 'broadcast')
Why it's wrong here
There is no 'broadcast' join type in the Spark join API. While SQL supports hints like /*+ BROADCAST(df2) */, the PySpark DataFrame API requires the broadcast function to be called explicitly on the DataFrame object itself rather than passed as a join type argument.
- ✗
spark.conf.set('spark.sql.broadcastJoin', 'true')
Why it's wrong here
This configuration setting does not exist in Spark. While there are configurations related to broadcast thresholds (e.g., spark.sql.autoBroadcastJoinThreshold), you cannot globally enable or force broadcast joins via a simple boolean flag without considering the size limitations of the data being broadcasted.
- ✗
df1.join(hint('broadcast'), df2, 'key')
Why it's wrong here
The hint function exists in the Spark SQL API but cannot be used as a direct argument in the DataFrame join method in this manner. The correct way to apply a hint in the DataFrame API is using the .hint() method before the join or using the explicit broadcast function.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.