DP-203 Develop data processing Practice Question
Network Topology
Refer to the exhibit. You submit a Spark job in Azure Synapse Analytics using the Azure CLI. The job runs slowly during the shuffle phase. The input data is about 200 GB. Which configuration change would best improve performance for this shuffle-heavy workload?
⚠ Common exam trap
DP-203 often tests the misconception that adding executors or memory fixes shuffle slowness — the real lever is partition count, and candidates must recognize that the default 200 is almost always too low for large datasets.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase spark.sql.shuffle.partitions to 800.
Spark's shuffle phase is governed by spark.sql.shuffle.partitions, which defaults to 200. With 200 GB of input, 200 partitions means roughly 1 GB per partition, causing large shuffle blocks, disk spills, and long task runtimes. Raising it to 800 creates smaller, more parallel partitions that better utilize the cluster's cores and reduce per-task memory pressure, which is the standard tuning lever for shuffle-heavy jobs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of executors to 4.
Why it's wrong here
Four executors cannot parallelise a 200 GB shuffle across enough partitions, leaving each task handling too much data and spilling. Raising executor count is the right lever for small datasets needing modest concurrency, but here the bottleneck is per-task shuffle volume, not executor scarcity.
- ✗
Change executor size to 'Large' to increase memory per executor.
Why it's wrong here
Enlarging executor memory gives each task more heap but does not reduce the volume of data moved across the network during shuffle, which is the actual constraint. Larger executors suit memory-intensive caching or broadcast joins, not a shuffle bound by partition count and data movement.
- ✓
Increase spark.sql.shuffle.partitions to 800.
Why this is correct
Shuffle partitions default to 200, which is too few for 200 GB and leaves each task handling roughly 1 GB. Raising spark.sql.shuffle.partitions to 800 increases parallelism across the shuffle stage, reducing per-task data volume and skew.
- ✗
Decrease spark.sql.shuffle.partitions to 200 to reduce overhead.
Why it's wrong here
Lowering shuffle partitions to 200 concentrates 200 GB into fewer, larger partitions, increasing per-task data and spill rather than relieving the shuffle. Reducing partitions helps small datasets where many tiny tasks create scheduling overhead, the opposite of this workload's profile.
Go deeper
Related to this question
About these practice questions
One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.