Courseiva

Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question

A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?

⚠ Common exam trap

Candidates often choose metrics like 'Shuffle read size' or 'JVM memory usage', which are related to performance but do not directly expose the uneven distribution of data that characterizes skew.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Task duration

Data skew occurs when one partition is significantly larger than others, causing one task to take much longer than the rest. By examining task duration and task input size, an engineer can spot outliers where a single executor handles a disproportionate amount of data. Identifying these metrics allows the engineer to apply remediation techniques like salting or repartitioning to balance the workload across the cluster effectively.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Task duration

    Why this is correct

    Task duration highlights significant variations in processing time across different executors. If most tasks complete in seconds but a few take minutes, it is a clear indicator of data skew or resource contention, prompting further investigation into the specific data distribution of the underlying partition.

  • ✗

    Executor CPU usage

    Why it's wrong here

    High CPU usage is often a result of computation, but it is not a direct indicator of data skew. Skewed data causes uneven task execution times, but CPU metrics might show high utilization across all nodes due to complex logic or serialized operations unrelated to partition distribution.

  • ✓

    Task input size

    Why this is correct

    Task input size directly reveals if specific partitions are much larger than the average. When one task processes gigabytes while others process megabytes, the imbalance is clearly data skew. This metric is the most reliable way to quantify the distribution of data across the cluster partitions.

  • ✗

    Shuffle read size

    Why it's wrong here

    Shuffle read size indicates the volume of data moved across the network during operations like joins or aggregations. While relevant to overall performance, it does not isolate the skew occurring within a specific stage as effectively as task-level input metrics do for identifyng partition imbalances.

  • ✗

    Driver memory usage

    Why it's wrong here

    Driver memory usage tracks the state of the master node. Skewed data affects the worker nodes performing the actual tasks, not the driver's memory consumption. Therefore, looking at driver memory is misleading when troubleshooting issues related to uneven data distribution across the distributed Spark executors.

About these practice questions

Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.