Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A streaming DataFrame job on Databricks writes to a Delta table every 10 seconds. Over several hours, the number of files in the target directory grows into the hundreds of thousands, and downstream reads slow dramatically. The job uses foreachBatch with a write that produces many small files per micro-batch. Which action should be taken to reduce the small-files problem for this streaming write?
⚠ Common exam trap
The trap here is reaching for repartition(1) or a single shuffle partition to reduce file count, which fixes file count by destroying write parallelism.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the write with optimizeWrite enabled (or set the target table property delta.autoOptimize.optimizeWrite = true) so Spark coalesces output files per partition before committing.
The streaming job creates many small files because each micro-batch writes multiple files per partition. Optimize Write coalesces those files into fewer, larger ones at commit time, directly reducing file proliferation and improving downstream read performance without sacrificing parallelism or increasing latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Call repartition(1) on the streaming DataFrame before the write in foreachBatch.
Why it's wrong here
Repartitioning to a single partition funnels all records through one task, bottlenecking the write and risking OOM. While it would produce fewer files, the throughput penalty and lack of scalability make it a poor solution compared with a write-time coalescing feature designed for this purpose.
- ✓
Configure the write with optimizeWrite enabled (or set the target table property delta.autoOptimize.optimizeWrite = true) so Spark coalesces output files per partition before committing.
Why this is correct
Optimize Write coalesces the many small files produced by each micro-batch into fewer, larger files before they are committed to the Delta table. This directly attacks file proliferation at write time, reducing the file count and improving downstream read performance without changing the streaming logic.
- ✗
Increase the trigger interval from 10 seconds to 10 minutes so each micro-batch writes more data.
Why it's wrong here
A longer trigger reduces the frequency of commits but does not coalesce files within a batch, and it increases latency. Each micro-batch may still emit many small files depending on partitioning, so the file count can continue to grow, just less often, without solving the underlying write-side fragmentation.
- ✗
Set spark.sql.shuffle.partitions to 1 so all output is written by a single task.
Why it's wrong here
Forcing one shuffle partition serializes the entire write through a single task, which severely limits throughput and can cause memory pressure on one executor. It may reduce file count but at an unacceptable performance cost and is not a targeted fix for streaming small files.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.