Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst is reviewing the Query Profile in Databricks SQL for a query that spills data to disk during a sort operation. Which two metrics should the analyst examine to confirm and diagnose the spill? (Choose two.)
⚠ Common exam trap
A common mix-up: candidates confuse shuffle-related metrics with spill metrics; spill is a distinct phenomenon that can occur without shuffles and is measured by specific spill counters.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Spill (Disk) Size
Spill (Disk) Size directly quantifies data written to disk due to memory overflow, confirming spill. Peak Execution Memory indicates memory pressure that causes spill. Together, they diagnose spill during sort. Other metrics like shuffle read size or output rows do not directly measure spill and are less relevant for this specific issue.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Spill (Disk) Size
Why this is correct
Spill (Disk) Size is a metric in the Query Profile that specifically reports the amount of data written to disk due to memory overflow during operations like sort or aggregation. A non-zero value confirms that spilling occurred. The analyst should examine this metric to verify the spill and its magnitude, which helps in tuning memory configurations or query logic.
- ✗
Number of Output Rows
Why it's wrong here
Number of Output Rows indicates the row count of the query result, which is not directly related to spilling. A high row count might suggest a large dataset, but it does not confirm spill. Spill is about data size relative to memory, not just row count. Therefore, this metric is not useful for diagnosing spill.
- ✗
Total Time Spent in Shuffle
Why it's wrong here
Total Time Spent in Shuffle measures the time taken for shuffle operations, which can be high in queries with large shuffles. However, it does not indicate spill during a sort. Spill can occur even without shuffles, such as in a global sort. Thus, this metric is not specific to spill diagnosis and should not be used to confirm it.
- ✗
Shuffle Read Size
Why it's wrong here
Shuffle Read Size indicates the amount of data read during a shuffle, which can be large if the query involves a shuffle. However, it does not directly indicate disk spill during a sort. Spill is a separate phenomenon where the sort operator exceeds memory and writes intermediate data to disk. While a large shuffle read might contribute to memory pressure, it is not a direct metric for spill.
- ✓
Peak Execution Memory
Why this is correct
Peak Execution Memory shows the maximum memory used by the query's execution. If this value approaches or exceeds the available memory per executor, it indicates memory pressure that can lead to spilling. By comparing Peak Execution Memory to the configured memory limits, the analyst can diagnose whether insufficient memory caused the spill during the sort operation.
About these practice questions
Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.