Courseiva
Managing Data →mediumMultiple Choice

Databricks-DA-Assoc Managing Data Practice Question

An analyst has a Delta table `bronze.raw_events` that has accumulated many small files over months of streaming writes. They now need to optimize read performance for downstream dashboards. Which Databricks SQL command should they run?

⚠ Common exam trap

Many candidates confuse VACUUM, which deletes unreferenced old files, with OPTIMIZE, which compacts current small files into larger ones.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

`OPTIMIZE bronze.raw_events;`

The small-file problem from streaming writes is resolved by compaction, and OPTIMIZE performs exactly that by merging small files into larger ones. Statistics gathering and vacuuming address different concerns, and write-time properties only influence future writes rather than existing layout. Compaction is the correct maintenance action for improving scan performance on a fragmented Delta table.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    `VACUUM bronze.raw_events;`

    Why it's wrong here

    VACUUM removes old data files that are no longer referenced by the transaction log and are past the retention threshold. It reclaims storage but does not compact small files into larger ones, so it will not improve read performance for the dashboards. Running VACUUM also reduces time-travel history, which is an unintended side effect unrelated to the stated goal.

  • ✗

    `ANALYZE TABLE bronze.raw_events COMPUTE STATISTICS;`

    Why it's wrong here

    Computing statistics collects metadata about data distribution used by the query optimizer to produce better plans. While useful, it does not change the physical layout of files on storage, so the small-file problem remains and scan overhead persists. Statistics help planning, not the underlying I/O cost of reading thousands of tiny files, so this does not solve the stated performance issue.

  • ✓

    `OPTIMIZE bronze.raw_events;`

    Why this is correct

    OPTIMIZE compacts many small files into fewer, larger files, which reduces file-open overhead and improves scan performance for dashboards. This directly addresses the small-file problem created by streaming writes. It is the standard maintenance command for this scenario and does not require changing the table definition or rewriting queries against the table.

  • ✗

    `ALTER TABLE bronze.raw_events SET TBLPROPERTIES ('delta.autoOptimize.optimizeWrite' = 'true');`

    Why it's wrong here

    Setting optimizeWrite affects future writes by reducing the number of files produced per write, but it does not compact the existing small files already present in the table. The dashboards query current data, so historical fragmentation still hurts performance. This property would help going forward but cannot remediate the accumulated small-file layout that already exists.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.