DP-203 Develop data processing Practice Question
You are optimizing a Spark DataFrame transformation in Azure Synapse Analytics. The DataFrame has 20 columns and 100 million rows. You notice that the job is slow due to many small files being written to the output. Which two actions can you take to reduce the number of output files? (Choose two.)
⚠ Common exam trap
Watch out — candidates often confuse `coalesce()` with `repartition()`, assuming both cause a shuffle, or they mistakenly think increasing partitions (Option D) will improve performance when it actually exacerbates the small-file issue.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use coalesce() to reduce the number of partitions without a shuffle.
Option A is correct because coalesce() reduces the number of partitions by merging existing partitions without performing a full shuffle, which directly decreases the number of output files written and is efficient when reducing partitions. Option E is correct because repartition() with a smaller number of partitions also reduces the partition count, and although it triggers a shuffle, it results in fewer output files being written. Option B is incorrect because caching only stores the DataFrame in memory or disk to speed up repeated access; it does not change the number of partitions or output files. Option C is incorrect because bucketing organizes data into buckets for join or query optimization and does not reduce the number of files written by a DataFrame write operation. Option D is incorrect because increasing partitions with repartition() using a larger number would create more partitions and therefore more output files, worsening the small-file problem.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use coalesce() to reduce the number of partitions without a shuffle.
Why this is correct
coalesce() merges partitions into the requested count without a full shuffle, so each task writes fewer, larger files. On a 100-million-row DataFrame this directly reduces the small-file problem while avoiding the network cost of repartitioning.
- ✗
Enable caching on the DataFrame before writing.
Why it's wrong here
Caching stores the DataFrame in memory to avoid recomputation across repeated actions; it has no effect on how many files the write operation emits. It tempts because caching speeds up iterative Spark jobs, but the small-file symptom stems from partition count, not recomputation.
- ✗
Apply bucketing on a column to group data.
Why it's wrong here
Bucketing controls data layout within partitions for join and shuffle efficiency; it does not coalesce the many small files written per partition. It tempts because it is a physical data-organisation technique, but it is the right choice for repeated join keys, not for consolidating output file counts.
- ✗
Increase the number of partitions using repartition() with a larger number.
Why it's wrong here
Increasing partition count multiplies the number of output files, worsening the small-file problem rather than reducing it. Repartitioning upward is tempting because it raises write parallelism, which helps when partitions are too few and each task is overloaded, not when many tiny files are produced.
- ✓
Use repartition() with a smaller number of partitions.
Why this is correct
repartition() with a smaller partition count performs a full shuffle that redistributes rows evenly, so each output task writes one larger file. This reduces the number of small files, though it costs more than coalesce() because of the shuffle.
Go deeper
Related to this question
About these practice questions
This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.