Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

A developer runs a PySpark job that joins a 10 GB DataFrame with a 50 MB lookup DataFrame. The job takes far longer than expected, and the Spark UI shows a SortMergeJoin with a large shuffle read and write for both sides. The developer wants the smallest change that most improves performance. Which action should the developer take?

⚠ Common exam trap

The trap here is tuning shuffle partitions or caching when the join strategy itself is the problem and a broadcast would eliminate the shuffle entirely.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Broadcast the 50 MB lookup DataFrame using broadcast() or by raising spark.sql.autoBroadcastJoinThreshold above 50 MB.

The Spark UI shows a SortMergeJoin with large shuffle on both sides, which is unnecessary when one input is only 50 MB. Broadcasting the small lookup table turns the join into a BroadcastHashJoin, removes the shuffle and sort of both DataFrames, and requires only a hint or a threshold adjustment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase spark.sql.shuffle.partitions so the sort-merge join shuffle uses more tasks and completes faster.

    Why it's wrong here

    More shuffle partitions can reduce per-task work, but the job still pays the full cost of shuffling and sorting both sides of the join. Since the lookup table is only 50 MB, avoiding the shuffle entirely via broadcast is far more effective than tuning partition counts for an unnecessary shuffle.

  • ✓

    Broadcast the 50 MB lookup DataFrame using broadcast() or by raising spark.sql.autoBroadcastJoinThreshold above 50 MB.

    Why this is correct

    Broadcasting the small side sends a copy of the 50 MB lookup to every executor and converts the join to a BroadcastHashJoin, eliminating the shuffle of both DataFrames. This is the minimal change that removes the expensive sort-merge shuffle and typically yields the largest speedup for a small-to-large join.

  • ✗

    Repartition the 50 MB lookup DataFrame by the join key before the join to co-locate matching rows.

    Why it's wrong here

    Repartitioning the small DataFrame still requires a shuffle of the large side to align keys, so a sort-merge shuffle remains. It adds an extra shuffle for the small table without eliminating the dominant cost, whereas broadcasting avoids shuffling both sides entirely.

  • ✗

    Cache the 10 GB DataFrame with persist(StorageLevel.MEMORY_AND_DISK) before the join.

    Why it's wrong here

    Caching the large side may help if it is reused across multiple actions, but it does not change the join strategy and does not remove the shuffle. For a single join, the cache consumes memory and disk without addressing the sort-merge shuffle that dominates runtime.

About these practice questions

This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.