Courseiva
Develop data processing →hardMultiple Choice

DP-203 Develop data processing Practice Question

Network Topology
name MyJobfile abfss://container@storage.dfs.core.windows.net/path/etl.pyexecutor-size Smallexecutors 2conf spark.sql.shuffle.partitions=400"

Refer to the exhibit. You submit a Spark job in Azure Synapse Analytics using the Azure CLI. The job runs slowly during the shuffle phase. The input data is about 200 GB. Which configuration change would best improve performance for this shuffle-heavy workload?

⚠ Common exam trap

DP-203 often tests the misconception that adding executors or memory fixes shuffle slowness — the real lever is partition count, and candidates must recognize that the default 200 is almost always too low for large datasets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase spark.sql.shuffle.partitions to 800.

Spark's shuffle phase is governed by spark.sql.shuffle.partitions, which defaults to 200. With 200 GB of input, 200 partitions means roughly 1 GB per partition, causing large shuffle blocks, disk spills, and long task runtimes. Raising it to 800 creates smaller, more parallel partitions that better utilize the cluster's cores and reduce per-task memory pressure, which is the standard tuning lever for shuffle-heavy jobs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of executors to 4.

    Why it's wrong here

    Four executors cannot parallelise a 200 GB shuffle across enough partitions, leaving each task handling too much data and spilling. Raising executor count is the right lever for small datasets needing modest concurrency, but here the bottleneck is per-task shuffle volume, not executor scarcity.

  • ✗

    Change executor size to 'Large' to increase memory per executor.

    Why it's wrong here

    Enlarging executor memory gives each task more heap but does not reduce the volume of data moved across the network during shuffle, which is the actual constraint. Larger executors suit memory-intensive caching or broadcast joins, not a shuffle bound by partition count and data movement.

  • ✓

    Increase spark.sql.shuffle.partitions to 800.

    Why this is correct

    Shuffle partitions default to 200, which is too few for 200 GB and leaves each task handling roughly 1 GB. Raising spark.sql.shuffle.partitions to 800 increases parallelism across the shuffle stage, reducing per-task data volume and skew.

  • ✗

    Decrease spark.sql.shuffle.partitions to 200 to reduce overhead.

    Why it's wrong here

    Lowering shuffle partitions to 200 concentrates 200 GB into fewer, larger partitions, increasing per-task data and spill rather than relieving the shuffle. Reducing partitions helps small datasets where many tiny tasks create scheduling overhead, the opposite of this workload's profile.

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.