Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question
A data engineer notices that a Databricks job processing a Delta table with 10,000 partitions runs slowly. The job filters on a column that is not the partition column, and the query plan shows that all partitions are being scanned. The engineer wants to improve performance without repartitioning the table. Which feature should be used?
⚠ Common exam trap
The trap here is assuming that partitioning is the only way to achieve data skipping, overlooking Z-ORDER BY as a complementary technique for non-partition columns.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Z-ORDER BY on the filter column during OPTIMIZE
When a filter is on a non-partition column, Delta Lake can still skip files if it has min/max statistics for that column and the data is clustered. Z-ORDER BY reorganizes data so that related values are stored together, improving the effectiveness of data skipping. Running OPTIMIZE with Z-ORDER BY on the filter column allows the query to read only the relevant files, avoiding a full scan of all partitions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enabling Change Data Feed on the table
Why it's wrong here
Change Data Feed records row-level changes for downstream consumers. It does not affect how data is laid out or how queries skip files. Enabling it adds metadata and storage overhead but does nothing to improve filter performance on a non-partition column. It is unrelated to the problem of scanning all partitions.
- ✗
Setting delta.dataSkippingNumIndexedCols to 0
Why it's wrong here
This property controls how many leading columns are indexed for data skipping. Setting it to 0 disables data skipping entirely, which would make performance worse, not better. The engineer needs to enable effective skipping on the filter column, not disable the feature. This option is the opposite of the required action.
- ✗
Increasing the number of shuffle partitions
Why it's wrong here
The spark.sql.shuffle.partitions setting controls the number of partitions used during shuffles such as joins and aggregations. It does not influence file skipping or partition pruning for a filter on a non-partition column. Increasing it may even add overhead if the data volume does not justify more partitions, and it does not address the full scan.
- ✓
Z-ORDER BY on the filter column during OPTIMIZE
Why this is correct
Z-ORDER BY co-locates related data in the same set of files, allowing Delta Lake to skip files based on min/max statistics for the Z-ordered column. When the filter column is Z-ordered, the query can skip many files even if it is not the partition column, dramatically reducing the amount of data scanned and improving performance without changing the partitioning scheme.
About these practice questions
One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.