Courseiva
Analyzing Queries →hardMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

An analyst is troubleshooting a query that hangs indefinitely during a join. Which TWO metrics in the Query Profile should the analyst examine to diagnose the issue?

⚠ Common exam trap

Candidates often look at general cluster metrics like total memory instead of join-specific indicators such as shuffle bytes and row disparities.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Shuffle Read/Write bytes.

When a join hangs, it is often due to massive data movement (shuffle) or extreme data skew. Monitoring the shuffle size helps determine if the network or I/O is saturated. Checking the 'Max vs Median' row count helps identify skew. These metrics are the most reliable indicators of join failure, enabling analysts to decide whether to adjust join keys, use broadcast hints, or address underlying data distribution issues.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Shuffle Read/Write bytes.

    Why this is correct

    High shuffle bytes indicate that large amounts of data are being exchanged across the network nodes. If a query hangs, it may be due to this massive data movement exceeding the cluster's network or I/O bandwidth, suggesting a need to reduce the data volume or optimize the join strategy.

  • ✓

    Maximum vs Median rows per task.

    Why this is correct

    A large discrepancy between the maximum and median rows processed per task confirms data skew. This means one task is handling much more data than others, causing it to 'hang' while other tasks complete quickly. Identifying this allows for the implementation of salting or other skew-handling techniques in the SQL.

  • ✗

    The number of users logged into the workspace.

    Why it's wrong here

    The number of active users has no direct impact on the performance of a single query's execution plan. While workspace-wide concurrency might affect resource allocation if the warehouse is over-subscribed, it is not a metric that explains why a specific individual query is hanging during its join operation.

  • ✗

    The total number of tables in the schema.

    Why it's wrong here

    The number of tables in a schema is metadata that does not impact the execution plan of a specific query. The SQL engine only interacts with the objects explicitly referenced in the query, so the schema's total object count is irrelevant to identifying the bottleneck of an active join.

  • ✗

    The total size of the database logs.

    Why it's wrong here

    Database or transaction logs track operations for ACID compliance and recovery. They are not involved in query execution or performance tuning. Monitoring log sizes is a task for platform administrators, not for analysts trying to debug a query that is hanging due to join logic or data distribution.

About these practice questions

Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.