Courseiva

Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question

A data engineer is optimizing a Databricks job that performs a join between a large fact table and a small dimension table. The job is slow, and the engineer suspects that the join strategy is not optimal. Which TWO actions should the engineer take to improve performance? (Choose two.)

⚠ Common exam trap

The trap here is assuming that increasing shuffle partitions or caching always helps, when the real issue is the join strategy and the solution is to avoid the shuffle altogether.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable Adaptive Query Execution (AQE) by setting spark.sql.adaptive.enabled to true

Broadcasting the small dimension table avoids shuffling the large fact table, which is a major performance win. Enabling Adaptive Query Execution allows Spark to automatically choose a broadcast join at runtime if the dimension table is small enough, and to optimize other aspects of the query. Together, these actions address the inefficient join strategy and improve performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable Adaptive Query Execution (AQE) by setting spark.sql.adaptive.enabled to true

    Why this is correct

    AQE can automatically convert a sort-merge join to a broadcast join at runtime if it detects that one side is small enough. It also optimizes shuffle partitions and handles skew. Enabling AQE allows Spark to adapt the join strategy based on actual data sizes, which can improve performance without manual hints. This is a recommended best practice in Databricks.

  • ✓

    Broadcast the small dimension table using a broadcast hint

    Why this is correct

    Broadcasting the small dimension table sends a copy to each executor, allowing the join to be performed without shuffling the large fact table. This avoids a costly shuffle and can significantly speed up the join. Using a broadcast hint explicitly tells Spark to use this strategy, which is appropriate when one side is small enough to fit in memory.

  • ✗

    Cache the large fact table in memory

    Why it's wrong here

    Caching the large fact table may improve performance if it is reused multiple times, but it does not change the join strategy. If the join is a sort-merge join, caching will not eliminate the shuffle. Moreover, caching a large table can consume significant memory and cause eviction. It is not a targeted fix for an inefficient join strategy.

  • ✗

    Increase spark.sql.shuffle.partitions to a very high value

    Why it's wrong here

    Increasing shuffle partitions can help with large shuffles by creating more, smaller partitions, but if the join is broadcast, there is no shuffle at all. Setting it too high creates many small tasks and overhead. For a broadcast join, this setting is irrelevant. For a sort-merge join, it might help, but it is not a targeted fix for an inefficient join strategy.

  • ✗

    Repartition the large fact table by the join key before the join

    Why it's wrong here

    Repartitioning the large fact table by the join key forces a full shuffle, which is expensive and often unnecessary if a broadcast join is possible. It may help if both sides are large, but in this scenario the dimension table is small, so broadcasting is a better strategy. Repartitioning adds overhead and does not address the root cause of the slow join.

About these practice questions

This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.