Databricks-DA-Assoc Managing Data Practice Question
A data analyst is preparing a Delta table in Unity Catalog for a dashboard that must return results quickly and consistently. The analyst needs to reduce the number of small files and improve data skipping on a frequently filtered column. (Choose two.)
⚠ Common exam trap
The trap here is treating VACUUM or cluster scaling as performance tuning for file layout, when VACUUM only deletes old files and scaling does not change physical data organization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run OPTIMIZE on the table to compact small files into larger ones.
OPTIMIZE compacts small files into larger ones, reducing read overhead, while ZORDER BY co-locates related values so Delta can skip files using column statistics. Together they address both the small-file problem and the data-skipping requirement. VACUUM only deletes unreferenced files, views do not change storage layout, and larger clusters do not fix inefficient file organization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the SQL warehouse cluster size so more workers can scan small files in parallel.
Why it's wrong here
Adding workers increases parallelism but does not fix the underlying small-file inefficiency or improve data skipping. Each worker still pays per-file overhead, and the table layout remains unoptimized. Scaling compute masks the symptom while increasing cost, whereas compacting files and clustering data address the root cause.
- ✓
Run OPTIMIZE on the table to compact small files into larger ones.
Why this is correct
OPTIMIZE rewrites many small files into fewer, larger files, which reduces per-file overhead during reads. For a dashboard that scans the table repeatedly, fewer files mean less metadata and I/O work. This directly addresses the small-file problem described in the scenario and is a standard performance maintenance operation for Delta tables.
- ✓
Run ZORDER BY on the frequently filtered column to co-locate related data.
Why this is correct
ZORDER BY reorganizes data within files so that rows with similar values in the specified column are stored together. Combined with Delta's file-level statistics, this enables data skipping so queries filtering on that column read fewer files. The scenario explicitly asks to improve data skipping on a frequently filtered column, which is exactly what ZORDER BY provides.
- ✗
Convert the table to a view so queries always compute results from the latest data.
Why it's wrong here
A view is a saved query, not a storage optimization. Converting the table to a view does not reduce file count or add data-skipping statistics, and it may push computation to query time, hurting dashboard performance. This option confuses access patterns with physical layout tuning and does not satisfy either requirement.
- ✗
Run VACUUM with a retention of zero hours to remove old files and speed up reads.
Why it's wrong here
VACUUM removes files no longer referenced by the current table version, but a zero-hour retention can break time travel and concurrent readers. It does not compact small files or improve data skipping; it only deletes unreferenced data. Using zero retention risks data loss for readers still on older snapshots, so it is inappropriate for this performance goal.
About these practice questions
Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.