Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst runs a query on a Delta table with a WHERE clause on a timestamp column, but the query scans the entire table. The analyst checks the table's metadata and sees that the column is not a partition column but is frequently used in filters. The analyst wants to enable data skipping to avoid full scans. Which action should the analyst take?
⚠ Common exam trap
Many candidates confuse partitioning with Z-ordering; partitioning on a timestamp column creates many small partitions, while Z-ordering provides data skipping without partitioning overhead.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run OPTIMIZE with ZORDER on the timestamp column to co-locate related data and improve data skipping.
Z-ordering on a frequently filtered column clusters data so that min/max statistics can be used to skip files. This is ideal for high-cardinality columns like timestamps where partitioning would create too many small files. OPTIMIZE with ZORDER rearranges data without changing the table structure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set spark.sql.parquet.filterPushdown to true to enable filter pushdown on the timestamp column.
Why it's wrong here
Filter pushdown is already enabled by default in Spark and Delta Lake. It allows filters to be pushed down to the file scan, but without clustering or partitioning, it cannot skip files. The issue is not filter pushdown but lack of data skipping due to data layout. Z-ordering is needed to cluster data.
- ✓
Run OPTIMIZE with ZORDER on the timestamp column to co-locate related data and improve data skipping.
Why this is correct
Z-ordering on a frequently filtered column rearranges data files so that related values are clustered together. This allows Delta Lake to skip files using min/max statistics, even if the column is not a partition. Running OPTIMIZE with ZORDER on the timestamp column will improve data skipping for queries filtering on that column.
- ✗
Convert the table to a partitioned table using the timestamp column as the partition key.
Why it's wrong here
Partitioning on a timestamp column can lead to a large number of small partitions, which degrades performance. Also, partitioning is not recommended for high-cardinality columns like timestamps. Z-ordering is more suitable for data skipping on such columns without the overhead of partitioning.
- ✗
Enable Delta Lake change data feed on the table to track changes to the timestamp column.
Why it's wrong here
Change data feed tracks row-level changes for downstream consumption, not for improving query performance. It does not affect data skipping or file pruning. Enabling it would not help the query scan fewer files; it would only add metadata overhead.
About these practice questions
This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.