Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst runs a Databricks SQL query that joins a 40 GB Delta table to a 12 GB Delta table and the Query Profile shows two Exchange nodes surrounding a SortMergeJoin. The analyst wants to confirm whether the shuffle is the dominant cost before rewriting the query. Which Query Profile metric should the analyst inspect first to quantify the shuffle's impact?
⚠ Common exam trap
The trap here is assuming any single operator's row count explains join slowness, when the Exchange byte metrics are what actually size the shuffle.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Shuffle write bytes and shuffle read bytes on the Exchange nodes
The Exchange operators are the shuffle stages that redistribute rows by join key before the SortMergeJoin, so their shuffle write and read byte counts are the direct measurement of redistribution cost. Inspecting those byte totals tells the analyst whether the shuffle dominates runtime and justifies alternatives such as broadcast join or pre-partitioning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Peak memory usage of the driver process
Why it's wrong here
Driver memory reflects the memory consumed by the process coordinating the job, result collection, and plan handling, not the volume of data moved between executors. A shuffle-heavy SortMergeJoin can spill on executors while driver memory stays flat, so this metric does not isolate the Exchange cost the analyst is trying to measure.
- ✗
Rows read from the smaller table's scan operator
Why it's wrong here
Rows read reports how many rows the scan operator emitted for that table, which speaks to input volume rather than the cost of redistributing data across the cluster. A join can read few rows but still shuffle enormous intermediate data, so this metric alone cannot confirm that the Exchange phase is the dominant cost in the SortMergeJoin plan.
- ✗
Number of files scanned in the Delta table directory
Why it's wrong here
File counts describe the physical layout of the Delta table and affect scan parallelism and partition pruning, but they do not report the network and serialization cost of an Exchange. The analyst already sees two Exchange nodes in the plan, so file-level statistics will not quantify how much data those shuffle stages actually moved.
- ✓
Shuffle write bytes and shuffle read bytes on the Exchange nodes
Why this is correct
The Exchange nodes are the shuffle boundaries, so their shuffle write and read byte counts directly measure how much data Spark serialized, spilled, and pulled across the network. These values quantify the redistribution cost that dominates a SortMergeJoin with two Exchanges, letting the analyst confirm whether shuffle volume is the bottleneck before rewriting the query.
About these practice questions
Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.