Courseiva
Analyzing Queries →hardMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

A data analyst runs a query on a Delta table with a WHERE clause on a timestamp column, but the query scans the entire table. The analyst checks the table's metadata and sees that the column is not a partition column but is frequently used in filters. The analyst wants to enable data skipping to avoid full scans. Which action should the analyst take?

⚠ Common exam trap

Many candidates confuse partitioning with Z-ordering; partitioning on a timestamp column creates many small partitions, while Z-ordering provides data skipping without partitioning overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run OPTIMIZE with ZORDER on the timestamp column to co-locate related data and improve data skipping.

Z-ordering on a frequently filtered column clusters data so that min/max statistics can be used to skip files. This is ideal for high-cardinality columns like timestamps where partitioning would create too many small files. OPTIMIZE with ZORDER rearranges data without changing the table structure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set spark.sql.parquet.filterPushdown to true to enable filter pushdown on the timestamp column.

    Why it's wrong here

    Filter pushdown is already enabled by default in Spark and Delta Lake. It allows filters to be pushed down to the file scan, but without clustering or partitioning, it cannot skip files. The issue is not filter pushdown but lack of data skipping due to data layout. Z-ordering is needed to cluster data.

  • ✓

    Run OPTIMIZE with ZORDER on the timestamp column to co-locate related data and improve data skipping.

    Why this is correct

    Z-ordering on a frequently filtered column rearranges data files so that related values are clustered together. This allows Delta Lake to skip files using min/max statistics, even if the column is not a partition. Running OPTIMIZE with ZORDER on the timestamp column will improve data skipping for queries filtering on that column.

  • ✗

    Convert the table to a partitioned table using the timestamp column as the partition key.

    Why it's wrong here

    Partitioning on a timestamp column can lead to a large number of small partitions, which degrades performance. Also, partitioning is not recommended for high-cardinality columns like timestamps. Z-ordering is more suitable for data skipping on such columns without the overhead of partitioning.

  • ✗

    Enable Delta Lake change data feed on the table to track changes to the timestamp column.

    Why it's wrong here

    Change data feed tracks row-level changes for downstream consumption, not for improving query performance. It does not affect data skipping or file pruning. Enabling it would not help the query scan fewer files; it would only add metadata overhead.

About these practice questions

This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.