Databricks-DA-Assoc Analyzing Queries Practice Question
Exhibit
== Physical Plan ==
AdaptiveSparkPlan (isFinalPlan=true)
+- BroadcastHashJoin [id#12], [id#25], Inner, BuildRight
+- Filter (isnotnull(id#12) && (status#15 = 'ACTIVE'))
+- Scan parquet default.orders
+- BroadcastQueryStage 0
+- BroadcastExchange
+- Filter (isnotnull(id#25) && (category#30 = 'ELECTRONICS'))
+- Scan parquet default.productsRefer to the exhibit. An analyst examines the physical execution plan of a Spark SQL join query. Based on the provided plan output, which optimization strategy was automatically applied by Adaptive Query Execution (AQE)?
⚠ Common exam trap
Candidates often assume the optimizer always chooses the most efficient join method, failing to recognize the specific 'BroadcastHashJoin' signature in the plan as a result of runtime AQE optimizations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AQE dynamically converted a sort-merge join into a broadcast hash join based on runtime size statistics.
Adaptive Query Execution in Databricks dynamically optimizes query plans at runtime based on precise statistics gathered from completed shuffle stages. The exhibit clearly displays a BroadcastHashJoin where the right side of the join was converted from a sort-merge join into a broadcast join using a broadcast exchange stage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dynamic partition pruning was used to eliminate scanning partitions on the orders table.
Why it's wrong here
Dynamic partition pruning eliminates reading specific partitions of a table based on runtime filter values derived from the other side of a join. The plan shows a standard broadcast hash join on an ID column, not partition pruning on directory paths.
- ✓
AQE dynamically converted a sort-merge join into a broadcast hash join based on runtime size statistics.
Why this is correct
Adaptive Query Execution evaluates runtime size metrics after executing child stages. When the size of the filtered 'products' table fell below the broadcast threshold, AQE replaced the costly sort-merge join with an efficient broadcast hash join, avoiding a heavy shuffle.
- ✗
Skew join optimization was triggered to split the skewed partition keys into multiple sub-tasks.
Why it's wrong here
Skew join optimization splits excessively large partitions into smaller sub-tasks when severe data skew is detected during shuffle stages. The physical plan here utilizes a broadcast join mechanism rather than handling a skewed shuffle-hash or sort-merge join.
- ✗
The Catalyst optimizer pushed down aggregations past the join boundary to reduce intermediate data size.
Why it's wrong here
Pushing down aggregations past joins is a logical optimization rule that rearranges syntax trees during analysis. The exhibit depicts the physical execution plan showing scan, filter, and broadcast exchange operators rather than aggregate pushdown operations.
About these practice questions
Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.