Courseiva

Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question

A data engineer notices that a scheduled Delta Lake maintenance pipeline is running significantly slower than expected. Upon checking the table history, they see that hundreds of tiny, fragmented data files have accumulated due to frequent streaming micro-batches. Which specific optimization command should the data engineer run first to resolve this file-size bottleneck?

⚠ Common exam trap

Candidates often confuse OPTIMIZE with VACUUM, believing that cleaning up old versions will automatically combine active small files into larger ones, which is not what VACUUM does.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Invoke the OPTIMIZE table_name command to compact small files into larger, optimized data files and improve subsequent query execution efficiency.

Running OPTIMIZE reorganizes the layout of Delta Lake data by compacting small files into larger, uniform files of roughly 1 GB in size. This significantly reduces metadata overhead and improves scan performance for subsequent read queries. While VACUUM removes old physical files, it does not compact active files. OPTIMIZE directly addresses the core symptom of small file proliferation caused by frequent streaming micro-batches.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Run VACUUM table_name RETAIN 0 HOURS to instantly purge all historical file references and force file compaction across the entire dataset.

    Why it's wrong here

    The VACUUM command deletes historical files older than the retention threshold to reclaim storage space, but it does not compact active data files or improve scan performance for current queries. Running it with zero retention violates safety checks and fails to merge small files.

  • ✗

    Execute REFRESH TABLE table_name to clear the local metastore cache and force Spark to rebuild the underlying file index from scratch.

    Why it's wrong here

    REFRESH TABLE only invalidates cached metadata so Spark re-reads the file listing; it neither merges nor removes the small files, leaving the bottleneck intact. It is tempting when stale metadata is suspected, and would be correct after external file changes, not for compaction.

  • ✓

    Invoke the OPTIMIZE table_name command to compact small files into larger, optimized data files and improve subsequent query execution efficiency.

    Why this is correct

    OPTIMIZE compacts the hundreds of small fragmented files produced by frequent streaming micro-batches into larger, right-sized data files. This directly resolves the file-size bottleneck, reducing per-file overhead and improving subsequent query execution efficiency without altering the underlying data.

  • ✗

    Execute ALTER TABLE table_name SET TBLPROPERTIES (delta.autoOptimize.optimizeWrite = true) to rewrite past historical micro-batches retroactively.

    Why it's wrong here

    optimizeWrite only affects future writes; it cannot retroactively compact the hundreds of existing small files, so the bottleneck persists. It is tempting because it reduces file counts going forward, and would be correct as a preventive setting before streaming ingestion begins, not as remediation.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.