DP-203 Develop data processing Practice Question
You are developing an Azure Databricks notebook that reads a large Delta table, performs a join with a smaller reference table, and writes the result back to Delta Lake. The job runs on a cluster with autoscaling enabled and frequently spills to disk during the join. You need to reduce shuffle and improve performance without changing the result. Which action should you take?
⚠ Common exam trap
The trap here is reaching for shuffle-partition tuning or caching when the real fix for a large-to-small join is broadcasting the small side to eliminate the shuffle altogether.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Broadcast the smaller reference table in the join using a broadcast hint.
When joining a large Delta table with a small reference table, broadcasting the small table avoids shuffling the large table entirely. Each executor receives a copy of the small table and performs the join locally, which removes the network exchange and disk spills observed in the job. The result set is identical, and the memory cost is modest because the broadcast side is small.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Repartition the large Delta table by the join key before the join.
Why it's wrong here
Repartitioning by the join key triggers a full shuffle of the large table, which is exactly the expensive operation causing spills. Although it can co-locate matching keys, it does not reduce total shuffle volume compared with a broadcast join and adds an extra stage. For a large-to-small join, broadcasting the small side is the more effective optimization.
- ✗
Cache the large Delta table in memory before performing the join.
Why it's wrong here
Caching the large table consumes significant executor memory and, with autoscaling, can cause eviction and recomputation rather than eliminating the shuffle. The join still requires a shuffle unless one side is broadcast. Caching does not address the fundamental shuffle-and-spill problem for a large-to-small join and may worsen memory pressure on the cluster.
- ✓
Broadcast the smaller reference table in the join using a broadcast hint.
Why this is correct
Broadcasting the smaller table replicates it to every executor, eliminating the shuffle of the large Delta table during the join. This directly reduces disk spills and network I/O, which are the observed symptoms. Because the reference table is small, the memory overhead is acceptable, and the join result is unchanged, satisfying the requirement to improve performance without altering output.
- ✗
Increase the spark.sql.shuffle.partitions value from the default to a much larger number.
Why it's wrong here
Increasing shuffle partitions creates more, smaller partitions, which can reduce per-partition memory pressure but also increases scheduling overhead and does not eliminate the shuffle itself. For a join with a small table, the root cause is unnecessary shuffling of the large table, so this tuning only masks the symptom. It also risks producing many tiny output files in Delta Lake.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DP-203 question from scratch — 509 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.