Courseiva
Analyzing Queries →hardMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

Exhibit

== Physical Plan ==
AdaptiveSparkPlan (isFinalPlan=true)
+- SortMergeJoin [id#8], [id#22], Inner
   +- Sort [id#8 ASC NULLS FIRST], false, 0
      +- Exchange hashpartitioning(id#8, 200), [id#8]
         +- Scan parquet default.table_a
   +- Sort [id#22 ASC NULLS FIRST], false, 0
      +- Exchange hashpartitioning(id#22, 200), [id#22]
         +- Exchange hashpartitioning(id#22, 200), [id#22]
            +- Scan parquet default.table_a

Refer to the exhibit. A data analyst reviews a physical execution plan where a table is scanned and shuffled multiple times in succession during a self-join operation. What underlying anti-pattern in the query construction likely caused this redundant exchange?

⚠ Common exam trap

Candidates often blame external factors like cluster configuration for performance issues, overlooking the fundamental Spark behavior where un-cached DataFrames are recomputed from their source lineage upon every reference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The table was referenced multiple times in complex join logic without caching, causing Spark to recompute its lineage.

Redundant exchanges and scans in an execution plan often occur when a DataFrame or table is referenced multiple times in complex joins or transformations without being cached or persisted. Spark recomputes the lineage from storage for each reference, leading to duplicate scans and shuffles.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The analyst applied a broadcast hint to a table that exceeded the broadcast size configuration threshold.

    Why it's wrong here

    An over-threshold broadcast hint makes Catalyst fall back to a sort-merge shuffle once, not scan and shuffle the same table repeatedly. It is tempting because broadcast hints do govern join strategy, and would be correct when a small dimension table is being shuffled unnecessarily instead of broadcast.

  • ✓

    The table was referenced multiple times in complex join logic without caching, causing Spark to recompute its lineage.

    Why this is correct

    Repeated references to the same table without caching force Spark to re-read and re-shuffle its lineage for each join branch. Persisting or caching the table breaks this recomputation, eliminating the redundant exchanges visible in the physical plan.

  • ✗

    The partition column data type mismatch forced Catalyst to duplicate the partition directories.

    Why it's wrong here

    Partition column data types do not cause Catalyst to duplicate partition directories or re-shuffle a self-join repeatedly; a mismatch yields casting or scan pruning issues instead. It is tempting because partition pruning problems do appear in execution plans, but that scenario shows skipped partitions, not successive exchanges.

  • ✗

    Automatic table compaction via OPTIMIZE was running concurrently during the execution phase.

    Why it's wrong here

    OPTIMIZE compaction rewrites files asynchronously and does not inject extra Exchange operators into a query's physical plan. It is tempting because concurrent maintenance affects file layout and read performance, which would be relevant when diagnosing small-file overhead rather than redundant shuffles in a self-join.

About these practice questions

This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.