Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

A data engineer is using PySpark to process a large DataFrame and needs to reduce the number of partitions before writing to a Delta table to avoid creating too many small files. The DataFrame currently has 2000 partitions, each about 10 MB. The engineer wants to reduce the number of partitions to approximately 200 while minimizing data shuffling. Which approach is most appropriate?

⚠ Common exam trap

It's easy for candidates to confuse coalesce with repartition; coalesce avoids a full shuffle but may result in uneven partition sizes, while repartition ensures even distribution at the cost of a shuffle.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use coalesce(200) to reduce the number of partitions without a full shuffle.

The correct answer is to use coalesce(200). coalesce reduces the number of partitions by merging existing partitions without a full shuffle, which is efficient for decreasing partition count. It is particularly useful when the partitions are already reasonably sized and you just want to reduce the number of output files. repartition would work but causes a full shuffle, which is unnecessary here. The other options involve shuffles or are not designed for simple partition reduction.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use repartition(200) to evenly redistribute the data across 200 partitions.

    Why it's wrong here

    repartition performs a full shuffle of the data to evenly distribute it across the specified number of partitions. While it can achieve the desired partition count, it introduces significant network and disk I/O overhead. In this scenario, the engineer wants to minimize shuffling, so repartition is less efficient than coalesce. repartition is better when you need to increase partitions or achieve even distribution, but here the goal is to reduce partitions with minimal shuffle.

  • ✓

    Use coalesce(200) to reduce the number of partitions without a full shuffle.

    Why this is correct

    coalesce is designed to reduce the number of partitions by combining existing partitions without a full shuffle. It moves data from some partitions to others on the same executor, which is efficient for reducing partition count. In this scenario, coalesce(200) will merge the 2000 partitions into 200, each approximately 100 MB, without incurring a wide shuffle. This minimizes network overhead and is suitable when the goal is to reduce the number of output files.

  • ✗

    Use repartitionByRange(200) on a key column to control data distribution.

    Why it's wrong here

    repartitionByRange performs a range partitioning shuffle based on the values of a column, which is useful for sorting or range-based operations but involves a full shuffle. It does not minimize shuffling; rather, it introduces a shuffle to redistribute data by range. In this scenario, the objective is to reduce partitions with minimal shuffle, so this method is not appropriate. It would also require specifying a column, which is not mentioned as necessary.

  • ✗

    Use bucketBy(200, 'id') and saveAsTable to create 200 buckets.

    Why it's wrong here

    bucketBy is used for bucketing tables on a column, which can optimize joins and aggregations, but it requires a shuffle during the write and is not a method to simply reduce partitions. It also creates a bucketed table, which may not be desired if the target is a Delta table without bucketing. The scenario asks for reducing partitions before writing, not for bucketing. This approach would not achieve the goal efficiently and could introduce unnecessary complexity.

About these practice questions

Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.