A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?
OPTIMIZE compacts many small files into larger ones, reducing the number of tasks and improving read throughput. In this scenario, the Spark UI shows thousands of small files and many small tasks, which is a classic small-file problem. Compaction directly addresses this by merging files, leading to fewer, larger files that are more efficient to read.
Why this answer
The Spark UI reveals many small tasks and thousands of small files in the source Delta table, which is a small-file problem. OPTIMIZE compacts these small files into larger ones, reducing task overhead and improving read performance. Other options either do not address the root cause or could exacerbate the issue by creating more partitions.
Exam trap
The trap here is assuming that increasing shuffle partitions or enabling AQE will solve small-file problems, when in fact compaction is required.