Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst is investigating a query that uses a window function with PARTITION BY and ORDER BY, and the query is slow. The analyst suspects that data skew is causing some partitions to be much larger than others. Which approach should the analyst take to diagnose the skew in the Query Profile?
⚠ Common exam trap
The trap here is focusing on memory or spill metrics to diagnose skew, when the most direct indicator is the uneven distribution of data read during shuffle.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check the 'Shuffle Read Size' per task to see if some tasks read significantly more data than others.
Data skew in window functions is often revealed by uneven shuffle read sizes across tasks. The Query Profile provides task-level metrics, and a significant disparity in Shuffle Read Size indicates that some partitions are much larger, causing skew. This allows the analyst to confirm the skew and then apply mitigation techniques like salting or repartitioning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Check the 'Shuffle Read Size' per task to see if some tasks read significantly more data than others.
Why this is correct
In the Query Profile, the task-level metrics for a shuffle stage show the amount of data each task reads. If data is skewed, a few tasks will have much larger Shuffle Read Size than others, indicating that those partitions are handling more data. This is a direct way to identify skew in window functions, which often require shuffling data by the partition key.
- ✗
Look at the 'Peak Execution Memory' to identify tasks using more memory.
Why it's wrong here
Peak Execution Memory shows the maximum memory used by a task, which can be higher for skewed tasks. However, it is not a direct measure of data skew; other factors like inefficient algorithms can also cause high memory usage. To diagnose skew, the analyst needs to compare data volumes processed by tasks, not just memory usage.
- ✗
Review the 'Number of Output Rows' for each task to see if some produce more rows.
Why it's wrong here
Number of Output Rows per task can indicate skew if some tasks output many more rows than others. However, in window functions, the output rows are typically the same as input rows, and skew is about the distribution of data being processed, not necessarily the output. Moreover, this metric may not be available per task for all operators. Shuffle read size is a more direct measure.
- ✗
Examine the 'Spill (Disk) Size' to see if any tasks spilled to disk.
Why it's wrong here
Spill (Disk) Size indicates memory pressure during operations, but it does not directly show skew. While skewed tasks may spill due to larger data volumes, spill can occur without skew. Therefore, this metric alone is not a reliable indicator of skew. The analyst should look at data distribution metrics like shuffle read size to confirm skew.
About these practice questions
One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.