Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question
A data engineer runs a nightly Databricks job that reads a large Delta table and writes aggregated results to another Delta table. The cluster logs show many small files in the source table, and the job runtime has increased steadily over weeks. The engineer wants to reduce the number of files without rewriting the entire table. Which command should be used?
⚠ Common exam trap
Many candidates confuse file cleanup (VACUUM) or statistics collection (ANALYZE) with file compaction (OPTIMIZE), when only compaction directly reduces the number of small files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
OPTIMIZE sales
The small-file problem is a common cause of slow reads in Delta Lake. OPTIMIZE performs bin-packing to combine small files into larger ones, reducing the number of files that must be opened during a read. It can be run on the entire table or a subset using a WHERE clause, and it does not require rewriting the entire table, making it the appropriate maintenance command for this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
ANALYZE TABLE sales COMPUTE STATISTICS
Why it's wrong here
ANALYZE TABLE collects statistics used by the query optimizer to improve join ordering and filter estimation. It does not combine small files or change the physical layout of the data. While helpful for query planning, it will not reduce the number of files in the Delta table, so it fails to address the root cause of the slowdown.
- ✓
OPTIMIZE sales
Why this is correct
OPTIMIZE compacts small files into larger ones (bin-packing) and can be run on the entire table or with a WHERE clause on a subset. It does not require rewriting the whole table and is the standard Delta Lake maintenance command to address the small-file problem, directly improving read performance for the nightly aggregation job.
- ✗
ALTER TABLE sales SET TBLPROPERTIES ('delta.autoOptimize.optimizeWrite' = 'true')
Why it's wrong here
This table property enables optimized writes for future write operations, which can reduce the number of small files produced going forward. However, it does not compact the existing small files already present in the table. The engineer needs to address the current file layout, so this setting alone will not fix the existing performance problem.
- ✗
VACUUM sales RETAIN 0 HOURS
Why it's wrong here
VACUUM removes old, unreferenced files that are no longer part of the Delta transaction log. It does not merge small files into larger ones; it only deletes obsolete data files. Running VACUUM with zero retention is also dangerous because it can break time travel and concurrent readers. It does not solve the small-file read performance issue.
About these practice questions
One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.