Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst runs a Databricks SQL query on a Delta table that has 500,000 small files. The query applies a filter on a high-cardinality column but still scans all files, resulting in slow performance. The analyst wants to reduce the number of files scanned without changing the query logic. Which action should the analyst take?
⚠ Common exam trap
The trap here is assuming that increasing cluster resources or altering query syntax will solve a small-file problem, when the real fix is file compaction and statistics collection.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run OPTIMIZE on the table to compact small files and improve data skipping.
Compacting small files with OPTIMIZE reduces the number of files that must be opened and enables Delta data skipping through collected statistics. This directly addresses the root cause: too many small files causing excessive I/O. Other options either disable useful optimizations, misuse syntax, or only mask the problem without reducing file count.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set spark.sql.adaptive.enabled to false to force a static execution plan.
Why it's wrong here
Disabling adaptive query execution removes runtime optimizations such as coalescing partitions and dynamically switching join strategies. It does not reduce the number of files scanned and may worsen performance by preventing the engine from adapting to actual data sizes during execution.
- ✓
Run OPTIMIZE on the table to compact small files and improve data skipping.
Why this is correct
OPTIMIZE compacts small files into larger ones and collects statistics, which enables Delta data skipping. With 500,000 small files, the query must open each file; compaction reduces file count and allows the engine to skip files based on min/max statistics, directly improving scan performance without altering query logic.
- ✗
Increase the driver node size to handle the large number of files.
Why it's wrong here
Increasing driver size may help with metadata handling but does not reduce the number of files scanned by the query. The bottleneck is the sheer file count and lack of data skipping; a larger driver does not compact files or collect statistics, so performance remains poor.
- ✗
Add a Z-ORDER BY clause to the query to sort results by the filtered column.
Why it's wrong here
Z-ORDER BY is not a query clause; it is an option used with OPTIMIZE to co-locate related data. Adding it to a SELECT statement would cause a syntax error. Even if used correctly, Z-ORDER alone does not compact files; it must be combined with OPTIMIZE to be effective.
About these practice questions
One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.