Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

A Spark job writes a large DataFrame to a Delta table partitioned by date. The job is taking much longer than expected, and the Spark UI shows that many tasks are writing very small files. You have already set spark.sql.shuffle.partitions to 200. What is the most effective way to reduce the number of small files written?

⚠ Common exam trap

The trap here is thinking that increasing shuffle partitions or coalescing to one partition will solve small files, when optimized writes are the Databricks-specific feature designed for this.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable optimized writes by setting spark.databricks.delta.optimizeWrite.enabled to true.

Enabling optimized writes on Databricks triggers an automatic shuffle before writing to Delta, which coalesces data into fewer, larger files. This directly addresses the small file problem without sacrificing parallelism or requiring manual tuning. The other options either worsen the issue, introduce bottlenecks, or do not target the root cause of many small files from partitioned writes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set spark.sql.files.maxRecordsPerFile to a high value to combine records into fewer files.

    Why it's wrong here

    maxRecordsPerFile limits the number of records per file, but setting it high does not guarantee fewer files; it only caps the records per file. If the data is already partitioned into many small partitions, each partition still writes at least one file. This setting can help prevent files from becoming too large, but it does not address the root cause of many small files from many partitions.

  • ✗

    Use coalesce(1) before writing to reduce the number of output files.

    Why it's wrong here

    Coalesce(1) would force all data into a single partition, which eliminates small files but severely bottlenecks the write and may cause out-of-memory errors. For a large dataset, this is impractical and defeats parallelism. While it reduces file count, it does not scale and is not the most effective solution. A better approach is to use partitionBy with a reasonable number of partitions or use Delta's optimize features.

  • ✗

    Increase spark.sql.shuffle.partitions to 2000.

    Why it's wrong here

    Increasing shuffle partitions would create more partitions, potentially leading to even more small files written per partition. The problem is that each partition is writing many small files, likely due to too many partitions relative to the data size or due to dynamic partition overwrites. More partitions would exacerbate the small file issue, not reduce it. The goal is to coalesce data before writing.

  • ✓

    Enable optimized writes by setting spark.databricks.delta.optimizeWrite.enabled to true.

    Why this is correct

    Optimized writes automatically coalesce small files during the write operation by adding a shuffle step that reduces the number of output files based on the data size. This is specifically designed to mitigate the small file problem in Delta Lake on Databricks. It balances file sizes without manual repartitioning, making it the most effective solution for reducing small files while maintaining parallelism.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.