Courseiva
Analyzing Queries →mediumMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

Exhibit

{
  "query_plan": {
    "node": "SortMergeJoin",
    "children": [
      {"node": "Exchange", "partitioning": "hash(user_id)"},
      {"node": "Exchange", "partitioning": "hash(user_id)"}
    ]
  }
}

Refer to the exhibit. The execution plan shows a SortMergeJoin with two Exchange nodes. What is the primary cause of the performance impact here?

⚠ Common exam trap

Candidates often misinterpret Exchange nodes in execution plans as caching layers rather than recognizing them as heavy network shuffles during sort-merge joins.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The tables are being re-shuffled to align rows on the join key, causing network overhead.

The presence of two Exchange nodes indicates that the data is being repartitioned across the cluster to align join keys on the same worker nodes. This full shuffle is costly in terms of network I/O. Recognizing this pattern helps analysts understand why large joins can be slow and justifies investigating whether tables can be pre-partitioned to avoid these runtime exchanges.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The join keys are not sorted, forcing a full scan of both tables.

    Why it's wrong here

    SortMergeJoin handles unsorted keys by sorting them during the shuffle phase. The performance bottleneck is not the sorting itself, but the data movement required to bring the corresponding keys into the same partition across the network. The sort happens locally after the data is successfully exchanged.

  • ✓

    The tables are being re-shuffled to align rows on the join key, causing network overhead.

    Why this is correct

    Exchange nodes in a Spark execution plan represent a shuffle, which involves moving data over the network to ensure that rows with the same join key are processed on the same node. This is a heavy operation that occurs when tables are not already partitioned by the join column.

  • ✗

    The optimizer is unable to choose a broadcast join because both tables are too large.

    Why it's wrong here

    While true that both tables are likely large, the performance impact is specifically caused by the shuffle operation (the Exchange nodes). Simply identifying that broadcast is impossible does not explain the mechanical cost of the SortMergeJoin's dependency on the network-intensive data redistribution represented by the Exchange nodes.

  • ✗

    The cluster is running out of memory due to the high number of partitions.

    Why it's wrong here

    While memory pressure can be a secondary effect of large shuffles, the execution plan node itself does not inherently indicate an out-of-memory error. The primary performance characteristic shown here is the intentional re-partitioning of data across the network to facilitate a join between two distributed datasets.

About these practice questions

This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.