Courseiva
Analyzing Queries →mediumMultiple Select

Databricks-DA-Assoc Analyzing Queries Practice Question

A data analyst is reviewing a Databricks SQL query that runs slowly and opens the Query Profile. The analyst wants to identify whether an individual task is disproportionately slow compared with its peers, indicating skew. Which TWO areas of the Query Profile should the analyst examine? (Choose two.)

⚠ Common exam trap

The trap here is blaming cluster size for a slow stage, when skew is a data distribution problem that adding executors cannot solve.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The number of rows and bytes written by each shuffle partition

Skew manifests as an individual task running far longer than its peers because one partition holds disproportionate data. Task duration distribution within a stage exposes the straggler, while per-partition shuffle write rows and bytes show the imbalanced partition that causes it. Together these two views confirm skew rather than a general resource shortage.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The version of the Databricks Runtime used by the warehouse

    Why it's wrong here

    Runtime version determines available features and performance characteristics across the board, but it is not a diagnostic signal for a single straggling task. Two queries on the same runtime can differ entirely in skew behavior, so runtime version cannot explain why one task within a stage runs far longer than its peers.

  • ✓

    The number of rows and bytes written by each shuffle partition

    Why this is correct

    Per-partition shuffle write sizes reveal imbalance directly: if one partition receives vastly more bytes or rows than its peers, the downstream task processing it will run long. Comparing these counts across partitions confirms skew at the shuffle boundary and points to the join or grouping key that is unevenly distributed.

  • ✗

    The SQL warehouse cluster size configured for the query

    Why it's wrong here

    Warehouse cluster size is a configuration setting that affects overall parallelism and cost, not a per-task diagnostic. Increasing cluster size adds executors but does not fix skew, because the oversized partition still runs on a single task; the analyst needs task-level and partition-level evidence before deciding whether resizing helps.

  • ✗

    The total number of stages in the physical plan

    Why it's wrong here

    Stage count reflects how many shuffle boundaries the plan introduces, which speaks to overall plan structure rather than per-task imbalance. A query can have many stages with perfectly even task durations, and a skewed query can have few stages, so stage count alone cannot reveal that one task is disproportionately slow.

  • ✓

    The duration distribution of tasks within a single stage

    Why this is correct

    The task duration distribution for a stage shows whether one or a few tasks run far longer than the median. A long tail where a single task dominates the stage runtime is the classic signature of skewed data, because one partition holds many more rows than the others and becomes the bottleneck for the whole stage.

About these practice questions

One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.