Courseiva
Analyzing Queries →hardMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

A data analyst runs a query that joins a fact table to a dimension table and groups by a dimension attribute. The Query Profile shows a BroadcastHashJoin, yet the query still takes much longer than expected and one stage shows a single task consuming most of its runtime. The dimension table is small, so the analyst suspects the bottleneck is elsewhere. Which explanation best fits the evidence?

⚠ Common exam trap

The trap here is assuming the slow join is the culprit, when an efficient broadcast join can coexist with a skewed aggregation downstream.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The grouping key is skewed, so one task in the aggregation stage processes far more rows than the others

Because the broadcast join removes shuffle from the join itself, the remaining slow stage with one dominant task is the aggregation. An unevenly distributed grouping key concentrates most rows into one partition, so that task runs far longer than its peers. Addressing the skew in the group-by, not the join, is what improves runtime.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The grouping key is skewed, so one task in the aggregation stage processes far more rows than the others

    Why this is correct

    A BroadcastHashJoin eliminates shuffle for the join, so the remaining bottleneck is likely the post-join aggregation. If one dimension attribute value is extremely common, the group-by partitions unevenly and a single task handles most rows, producing exactly the long-running straggler task the profile shows despite the efficient join.

  • ✗

    The fact table lacks statistics, causing Spark to choose a nested loop join

    Why it's wrong here

    Spark does not fall back to a nested loop join for a standard equi-join in Databricks SQL, and the profile explicitly shows a BroadcastHashJoin, so join selection is not the problem. Missing statistics can influence plan choices, but they cannot explain a single task dominating a stage after an efficient broadcast join has already been chosen.

  • ✗

    The broadcast of the small table is failing and silently falling back to a sort-merge join

    Why it's wrong here

    A fallback from broadcast to sort-merge would appear in the plan as a SortMergeJoin with Exchange nodes, not as a BroadcastHashJoin. Since the profile already shows a BroadcastHashJoin, the broadcast succeeded, so the straggler task must originate from a different stage such as the aggregation rather than from a failed broadcast.

  • ✗

    The dimension table is too large to broadcast and must be repartitioned before joining

    Why it's wrong here

    The profile already shows a BroadcastHashJoin, which means Spark successfully broadcast the dimension table, so size was not a blocker here. Repartitioning a table that was already broadcast adds no benefit and does not address the single slow task, which points to data distribution in a later stage rather than to the join build side.

About these practice questions

This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.