Databricks-DA-Assoc Analyzing Queries Practice Question
An analyst is reviewing a slow-running query and wants to identify potential bottlenecks. Which TWO of the following metrics in the Databricks SQL query profile are most indicative of inefficient data distribution?
⚠ Common exam trap
Candidates often look only at total query execution time instead of specific distributed metrics like shuffle read/write bytes and skewed task durations to diagnose bottlenecks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Shuffle Read/Write bytes are significantly higher than the input data size.
Data skew and excessive shuffling are primary causes of performance degradation. High 'Shuffle Read' and 'Shuffle Write' values indicate large amounts of data moving across the network, while skewed task durations signal that one executor is doing significantly more work than others. Identifying these metrics early allows analysts to apply techniques like salting keys or partitioning to balance the workload across the cluster.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Shuffle Read/Write bytes are significantly higher than the input data size.
Why this is correct
High shuffle volumes compared to input data often suggest unnecessary data movement or poor join strategies. If the shuffle size is much larger than the original input, the query is likely re-distributing data inefficiently, which adds substantial network latency and increases the risk of executor memory errors.
- ✗
The 'Scan' node shows high throughput and low execution time.
Why it's wrong here
High throughput at the scan node is a positive indicator, suggesting that the data retrieval process is highly efficient. This indicates that the storage layer and file formats are performing well, and is not a symptom of a query bottleneck or an inefficient data distribution issue.
- ✓
Task duration distribution shows one task taking much longer than others.
Why this is correct
This is the classic symptom of data skew. When a single task processes a significantly larger partition than others, the entire job must wait for that single, overloaded task to finish, wasting cluster resources and delaying results even if other tasks complete almost instantaneously.
- ✗
The query shows a high number of file metadata operations.
Why it's wrong here
While many metadata operations can indicate a 'small file problem' that degrades scan performance, they are not direct indicators of inefficient data distribution during joins or aggregations. This metric identifies storage-level inefficiency rather than computational skew or unnecessary shuffle operations within the query's execution plan.
- ✗
The query plan indicates a broadcast hash join is being used.
Why it's wrong here
A broadcast hash join is generally an optimized join type that avoids shuffling. Seeing this in the query profile actually suggests that the optimizer has already successfully identified an efficient way to handle the join, rather than indicating an inefficiency that needs to be addressed by the analyst.
About these practice questions
This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.