Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A PySpark DataFrame job on Databricks runs slowly. Inspection of the Spark UI shows that a shuffle stage writes 200 partitions but downstream stages process only a few, and the physical plan shows an Exchange before a filter. Which two changes are most likely to improve performance? (Choose two.)
⚠ Common exam trap
The trap here is reaching for more shuffle partitions by reflex, when the UI actually shows too many near-empty partitions and excess shuffle input.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce the number of shuffle partitions or coalesce the output so fewer, larger partitions are produced.
The plan shows an Exchange feeding stages that process only a few partitions, meaning shuffle volume and partition sizing are both inefficient. Filtering earlier cuts the rows that ever reach the shuffle, and consolidating partitions removes the overhead of many tiny tasks. Together they reduce both the data moved across the network and the number of tasks scheduled, which is what the UI evidence points to. Resource or caching changes would not target either cause.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase spark.sql.shuffle.partitions to a much larger number so each task handles fewer rows.
Why it's wrong here
The symptom is that many output partitions are empty or nearly empty downstream, not that each task is too large. Raising the partition count would create even more tiny tasks and more scheduling overhead. The real issue is the volume of data entering the shuffle, which partitioning cannot reduce. This change would likely make the job slower.
- ✗
Repartition the DataFrame on the join key before the shuffle to colocate matching rows.
Why it's wrong here
An explicit repartition on the join key introduces an extra full shuffle before the join's own Exchange, doubling network and disk cost. Spark already partitions on the join key during the Exchange, so this is redundant. It does not reduce the amount of data shuffled and adds a stage. This would worsen rather than improve the runtime.
- ✓
Reduce the number of shuffle partitions or coalesce the output so fewer, larger partitions are produced.
Why this is correct
When a shuffle produces many near-empty partitions, consolidating them reduces task launch overhead and improves per-task efficiency. Coalescing after the shuffle or lowering the partition count creates fewer, better-sized partitions for downstream stages. This matches the observed pattern of many partitions with little data. It addresses the scheduling waste visible in the UI.
- ✗
Cache the shuffled DataFrame with persist so downstream stages reuse the same partitions.
Why it's wrong here
Caching helps when the same DataFrame is reused across multiple actions, but the problem described is wasted shuffle volume in a single lineage, not repeated reads. Persisting the post-shuffle data would hold large partitions in memory with no reuse to amortize the cost. It also risks eviction and recomputation. Caching is not the right remedy for excessive shuffle input.
- ✓
Apply the filter before the join or aggregation that triggers the Exchange so less data is shuffled.
Why this is correct
Pushing a filter earlier reduces the number of rows entering the shuffle, so the Exchange writes and reads far less data. Catalyst can sometimes do this automatically through predicate pushdown, but only when the filter is expressible on the source or the plan allows it. Manually reordering transformations guarantees the reduction. This directly addresses the wasted shuffle work shown in the UI.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.