Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

A data engineer is using Delta Lake to manage a table that receives frequent updates and deletes. The engineer notices that query performance has degraded over time due to many small files. Which command should be used to optimize the table by compacting small files and improving query performance?

⚠ Common exam trap

Many exam-takers confuse VACUUM with OPTIMIZE; VACUUM removes old files but does not compact small files, while OPTIMIZE specifically addresses file compaction.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

OPTIMIZE table_name

The OPTIMIZE command is designed to compact small files into larger ones, improving read performance. It can also perform Z-Ordering to cluster data. VACUUM removes old files but does not compact, ANALYZE collects statistics, and autoOptimize.optimizeWrite prevents future small files but does not fix existing ones. Therefore, OPTIMIZE is the correct command to address the current small file issue.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    ANALYZE TABLE table_name COMPUTE STATISTICS

    Why it's wrong here

    ANALYZE TABLE computes statistics used by the query optimizer to improve query plans. It does not compact files or reduce the number of small files. While statistics can help with performance, they do not address the underlying file fragmentation issue. This command is complementary but not the solution for small file compaction.

  • ✓

    OPTIMIZE table_name

    Why this is correct

    The OPTIMIZE command compacts small files into larger ones, reducing the number of files and improving query performance. It also supports Z-Ordering for multi-dimensional clustering. This is the standard Delta Lake operation for file compaction. It is efficient and can be scheduled regularly. This command directly addresses the issue of many small files.

  • ✗

    VACUUM table_name

    Why it's wrong here

    VACUUM removes old, unreferenced files to save storage, but it does not compact small files. It is used for data retention and cleanup, not performance optimization. Running VACUUM will not reduce the number of small files that are still referenced by the table. Therefore, it does not solve the performance degradation caused by small files.

  • ✗

    ALTER TABLE table_name SET TBLPROPERTIES ('delta.autoOptimize.optimizeWrite' = 'true')

    Why it's wrong here

    Setting optimizeWrite to true enables automatic compaction during writes, which can prevent small files from being created in the future. However, it does not compact existing small files. To address the current degradation, an explicit OPTIMIZE is needed. This property is useful for ongoing maintenance but does not fix the existing problem immediately.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.