Databricks-DE-Assoc Data Transformation and Modeling Practice Question
A data engineer is working on a Delta Lake table that has accumulated millions of small files due to frequent streaming updates. This fragmentation has significantly degraded query performance. Which operation should the engineer execute to optimize file layout without altering table data?
⚠ Common exam trap
Candidates often choose the VACUUM command instead of OPTIMIZE, confusing file compaction for small files with the deletion of historical snapshot files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Invoke the OPTIMIZE command to compact small parquet files into larger files using bin-packing.
Running the OPTIMIZE command on a Delta table compacts many small files into larger, optimized files using the bin-packing algorithm. This directly addresses fragmentation issues caused by frequent streaming micro-batches, improving read efficiency and scan speeds. Data engineers frequently schedule optimization jobs as part of bronze-to-silver and silver-to-gold maintenance pipelines in production environments to maintain optimal query performance and reduce cloud storage request latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Execute REFRESH TABLE to update the file metadata cache across all cluster worker nodes.
Why it's wrong here
The REFRESH TABLE command only invalidates and reloads cached metadata for a specific table in Apache Spark memory. It does not perform any physical file compaction or rewrite operations to combine fragmented small files into larger ones.
- ✗
Run the VACUUM command with a retention threshold of zero hours to purge all historical data files.
Why it's wrong here
VACUUM with zero-hour retention deletes files no longer referenced by the transaction log, reclaiming storage but leaving the small-file layout untouched, so queries remain slow. It is tempting as a cleanup command, but VACUUM removes obsolete data; OPTIMIZE compacts existing small files into larger ones without altering table contents.
- ✓
Invoke the OPTIMIZE command to compact small parquet files into larger files using bin-packing.
Why this is correct
The OPTIMIZE command groups small data files into larger, more efficient files, dramatically accelerating read performance for downstream analytical queries. It operates safely concurrently with reads and writes on Delta tables without causing job failures.
- ✗
Execute ALTER TABLE SET TBLPROPERTIES to enable automatic file compaction on every write transaction.
Why it's wrong here
Enabling auto-compaction via table properties does not consolidate the millions of existing small files; it only affects future writes, so current query performance stays degraded. It is tempting because it sounds like ongoing maintenance, but OPTIMIZE (with compaction) is the operation that rewrites existing files into larger ones.
About these practice questions
Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.