Databricks-DE-Pro Developing Code (Python/SQL) Practice Question
A data engineer is optimizing a PySpark job that processes a large DataFrame and writes the result to a Delta table. The job currently uses repartition(100) before writing, but the output consists of many small files. The engineer wants to reduce the number of output files without shuffling the entire dataset again. Which approach should be used?
⚠ Common exam trap
The trap here is thinking that repartition is always better for reducing file count, ignoring the shuffle cost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use coalesce(10) before writing to reduce the number of partitions without a full shuffle.
coalesce is the correct choice because it reduces the number of partitions without a full shuffle, making it efficient for reducing output file count. repartition would cause an unnecessary shuffle, and changing shuffle partitions or using partitionBy does not achieve the goal of reducing files without a shuffle.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use coalesce(10) before writing to reduce the number of partitions without a full shuffle.
Why this is correct
coalesce reduces the number of partitions by combining existing partitions without a full shuffle. It is more efficient than repartition when reducing partitions because it avoids shuffling data across the network. Using coalesce(10) would combine the 100 partitions into 10, reducing the number of output files while minimizing data movement.
- ✗
Use partitionBy("column") to write the data into subdirectories, which reduces the number of files per directory.
Why it's wrong here
partitionBy creates subdirectories based on the column values, but it does not reduce the total number of files; it may even increase the number of files per partition if there are many unique values. The goal is to reduce the overall file count, not to organize by column. This approach does not address the small file problem directly.
- ✗
Set spark.sql.shuffle.partitions to 10 before writing.
Why it's wrong here
spark.sql.shuffle.partitions controls the number of partitions used after a shuffle for SQL operations, but it does not affect the number of partitions in an existing DataFrame. Changing this property would not reduce the output files from the current DataFrame without triggering a shuffle. It is not the correct approach for reducing file count on write.
- ✗
Use repartition(10) before writing to redistribute data evenly and reduce file count.
Why it's wrong here
repartition(10) would perform a full shuffle to redistribute data into 10 partitions, which is expensive and unnecessary if the goal is simply to reduce file count. While it would reduce the number of output files, it incurs a significant performance cost due to the shuffle. coalesce is preferred when reducing partitions without needing balanced data distribution.
About these practice questions
This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.