Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst is analyzing a slow-running Databricks SQL query. The analyst suspects that the query is suffering from excessive shuffling. Which two actions should the analyst take to reduce shuffle and improve performance? (Choose two.)
⚠ Common exam trap
The trap here is thinking that increasing shuffle partitions or caching always helps; instead, AQE and broadcast joins are targeted at reducing shuffle volume and overhead.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Adaptive Query Execution (AQE) to dynamically coalesce shuffle partitions.
Enabling AQE allows dynamic coalescing of shuffle partitions and skew handling, reducing shuffle overhead. Using broadcast joins for small tables avoids shuffling the large table entirely. Both are effective strategies to minimize shuffle and improve query performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Repartition the large table on a high-cardinality column before the join.
Why it's wrong here
Repartitioning on a high-cardinality column may not align with the join key and could cause an unnecessary shuffle. Repartitioning should be done on the join key to co-locate data, but if the join key is not high-cardinality, it may not help. In general, repartitioning adds a shuffle step, so it should be used judiciously.
- ✓
Enable Adaptive Query Execution (AQE) to dynamically coalesce shuffle partitions.
Why this is correct
AQE can dynamically coalesce small shuffle partitions into larger ones, reducing the number of tasks and overhead. It also can optimize skew joins. Enabling AQE (spark.sql.adaptive.enabled=true) is a recommended practice to reduce shuffle-related inefficiencies, especially when partition sizes are uneven.
- ✗
Increase the number of shuffle partitions by setting spark.sql.shuffle.partitions to a very high value.
Why it's wrong here
Increasing shuffle partitions creates more, smaller partitions, which can increase overhead and not reduce shuffle. The goal is to have partitions of optimal size (e.g., 100-200 MB). Setting it too high leads to many small tasks and scheduling overhead, worsening performance.
- ✓
Use broadcast joins for small tables to avoid shuffling the large table.
Why this is correct
Broadcasting a small table eliminates the need to shuffle the large table for the join. This reduces network I/O and disk I/O, speeding up the query. It is a key technique to avoid shuffles when one side of the join is small enough to fit in memory on each executor.
- ✗
Cache the large table in memory to avoid shuffling it in subsequent operations.
Why it's wrong here
Caching can speed up repeated access to the same data, but it does not reduce shuffle for a single query. Shuffle occurs during operations like joins and aggregations. Caching the large table would consume memory and not address the shuffle caused by the join. It might even cause memory pressure.
About these practice questions
This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.