Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question
A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?
⚠ Common exam trap
Candidates often choose metrics like 'Shuffle read size' or 'JVM memory usage', which are related to performance but do not directly expose the uneven distribution of data that characterizes skew.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Task duration
Data skew occurs when one partition is significantly larger than others, causing one task to take much longer than the rest. By examining task duration and task input size, an engineer can spot outliers where a single executor handles a disproportionate amount of data. Identifying these metrics allows the engineer to apply remediation techniques like salting or repartitioning to balance the workload across the cluster effectively.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Task duration
Why this is correct
Task duration highlights significant variations in processing time across different executors. If most tasks complete in seconds but a few take minutes, it is a clear indicator of data skew or resource contention, prompting further investigation into the specific data distribution of the underlying partition.
- ✗
Executor CPU usage
Why it's wrong here
High CPU usage is often a result of computation, but it is not a direct indicator of data skew. Skewed data causes uneven task execution times, but CPU metrics might show high utilization across all nodes due to complex logic or serialized operations unrelated to partition distribution.
- ✓
Task input size
Why this is correct
Task input size directly reveals if specific partitions are much larger than the average. When one task processes gigabytes while others process megabytes, the imbalance is clearly data skew. This metric is the most reliable way to quantify the distribution of data across the cluster partitions.
- ✗
Shuffle read size
Why it's wrong here
Shuffle read size indicates the volume of data moved across the network during operations like joins or aggregations. While relevant to overall performance, it does not isolate the skew occurring within a specific stage as effectively as task-level input metrics do for identifyng partition imbalances.
- ✗
Driver memory usage
Why it's wrong here
Driver memory usage tracks the state of the master node. Skewed data affects the worker nodes performing the actual tasks, not the driver's memory consumption. Therefore, looking at driver memory is misleading when troubleshooting issues related to uneven data distribution across the distributed Spark executors.
About these practice questions
Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.