Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst is investigating why a specific Spark SQL query takes an exceptionally long time to complete and exhibits signs of severe data skew. Which TWO metrics or behaviors in the Databricks Spark UI typically indicate that data skew is impacting the query? (Choose two)
⚠ Common exam trap
Students often mistake general cluster slowness for data skew, failing to check the specific percentile task duration gaps in the Spark UI.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The maximum task duration is drastically higher than the 75th and median task durations for a given stage.
Data skew occurs when records are unevenly distributed across partitions, causing a few tasks to process the vast majority of data while others finish quickly. In the Spark UI, this manifests as a massive disparity between task duration percentiles and extremely high memory consumption or spill metrics on the tasks handling the skewed keys.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The maximum task duration is drastically higher than the 75th and median task durations for a given stage.
Why this is correct
A wide gap between the median task duration and the maximum task duration is a primary hallmark of data skew. Most tasks complete rapidly, but the stage cannot finish until the few tasks processing the skewed partition keys finish executing their workloads.
- ✗
The total number of input files scanned is zero because all data is successfully pruned by partition filters.
Why it's wrong here
Zero input files scanned indicates successful partition pruning and metadata-only query execution or empty results, which is a sign of optimal data skipping rather than an indicator of workload performance degradation caused by data skew.
- ✓
Specific tasks show high amounts of spill to disk for memory-heavy operations like aggregations and joins.
Why this is correct
Skewed partitions concentrate disproportionately large volumes of data into single tasks. This accumulation overwhelms available executor memory, forcing the task to spill intermediate shuffle or aggregation data to local disk, which drastically slows down execution.
- ✗
The driver node runs out of memory during the initial parsing phase before any execution stages are generated.
Why it's wrong here
Driver out-of-memory errors during the initial parsing phase typically stem from collecting massive datasets to the driver via operations like 'collect()' or parsing extremely complex SQL abstract syntax trees, rather than runtime data skew across worker task executors.
- ✗
All tasks across all worker nodes complete in nearly identical timeframes with minimal variance.
Why it's wrong here
Identical task durations indicate an exceptionally well-balanced workload where data partitions are uniformly distributed across cluster executors. This represents the ideal execution state, which is the exact opposite of a performance bottleneck caused by data skew.
About these practice questions
This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.