Databricks-GenAI-Assoc Data Preparation Practice Question
A data engineer is preparing a large Delta table of conversation logs for embedding generation. The table has frequent small appends, and the engineer needs to reduce file fragmentation and improve read throughput before the embedding job runs. Which two actions should the engineer take? (Choose two.)
⚠ Common exam trap
Watch out — candidates often confuse VACUUM, which deletes unreferenced files, with OPTIMIZE, which rewrites live data into fewer files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable auto compaction and optimized writes on the Delta table so small writes are coalesced automatically.
File fragmentation from frequent appends is addressed by consolidating small files. OPTIMIZE with optional ZORDER compacts and clusters existing data, while auto compaction and optimized writes prevent new fragmentation from accumulating. Together they reduce the number of files the embedding job must open and can improve predicate pushdown, directly improving read throughput.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Run VACUUM with a retention of zero hours to delete old files and free space before the embedding job.
Why it's wrong here
VACUUM removes files no longer referenced by the transaction log; it does not merge the current small files into larger ones. A zero-hour retention also eliminates time travel history and can break concurrent readers, which is dangerous in production. Vacuuming reduces storage, not file count, so read throughput for the embedding job would not improve.
- ✗
Increase the number of shuffle partitions to the maximum supported value so that each task writes a single small file.
Why it's wrong here
More shuffle partitions create more output files, which worsens fragmentation rather than reducing it. Writing one small file per task is the opposite of compaction and increases metadata and listing overhead for downstream readers. This setting is useful for large shuffles, not for consolidating an append-heavy table before a read-heavy embedding job.
- ✓
Enable auto compaction and optimized writes on the Delta table so small writes are coalesced automatically.
Why this is correct
Auto compaction merges small files after a write, and optimized writes shuffle data so fewer, larger files are produced in the first place. Together they keep fragmentation low between manual maintenance runs, which suits a table receiving frequent small appends. This directly addresses the file-count problem before the embedding job reads the data.
- ✓
Run OPTIMIZE on the table to compact small files, optionally with ZORDER on columns frequently used in filters.
Why this is correct
OPTIMIZE rewrites many small files into fewer, larger files, which reduces per-file overhead during the embedding job's scans. Adding ZORDER on filter columns clusters related data so predicate pushdown skips more files. For an append-heavy table, periodic compaction is the standard Delta Lake remedy before a large read-heavy workload such as embedding generation.
- ✗
Convert the table to a Parquet directory and rely on the file system to merge small files during reads.
Why it's wrong here
Parquet directories do not merge files during reads, and converting away from Delta loses transaction logs, time travel, and the OPTIMIZE command itself. The file system lists and opens every small file, so read overhead persists. This change removes the very tooling that solves the fragmentation problem and adds no benefit for the embedding workload.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.