Databricks-DE-Pro Data Modelling Practice Question
Exhibit
{"table_name": "sales_data", "partition_columns": ["region", "date"], "z_order_columns": ["customer_id"], "file_format": "delta"}Refer to the exhibit. An engineer observes that queries filtering on 'customer_id' are running slowly despite Z-Ordering. What is the most likely cause?
⚠ Common exam trap
Candidates often blame the Z-Ordering configuration itself, assuming it is broken, rather than realizing that Z-Ordering cannot overcome the lack of partition pruning in the query filter.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The queries lack filters on 'region' or 'date', preventing effective partition pruning.
The exhibit shows that the table is partitioned by 'region' and 'date', while Z-Ordering is applied to 'customer_id'. If queries filter on 'customer_id' but do not provide 'region' or 'date', the engine must scan all partitions. Z-Ordering is only effective within each partition. If the data is not well-clustered or the partitions are too large, the engine cannot skip files effectively, leading to high latency during execution.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The partition columns should include 'customer_id' to improve the pruning speed.
Why it's wrong here
Adding high-cardinality columns like 'customer_id' to the partition strategy leads to the small file problem. This increases metadata overhead and creates thousands of unnecessary directories, which actually slows down query performance and makes file management significantly more difficult to automate within the Databricks environment.
- ✗
Z-Ordering must be performed on the partition columns instead of the join columns.
Why it's wrong here
Z-Ordering is meant for non-partition columns, typically those with high cardinality used in filters or joins. Z-Ordering on partition columns is redundant because partition pruning already handles the filtering logic. The issue here is the lack of partition pruning in the query, not the Z-Ordering configuration itself.
- ✓
The queries lack filters on 'region' or 'date', preventing effective partition pruning.
Why this is correct
Because the table is partitioned by region and date, failing to include these in the WHERE clause forces the engine to scan every partition. Z-Ordering only clusters data within individual partitions. If the engine doesn't prune the partitions first, the Z-Ordering benefits are largely ignored during the scan.
- ✗
The file format should be changed to Parquet to improve individual file read performance.
Why it's wrong here
Delta Lake is built on top of Parquet and provides features like Z-Ordering and partition pruning that raw Parquet does not support natively. Changing the format to Parquet would remove the ability to use Delta-specific optimizations, making the query performance even worse while losing critical ACID transaction capabilities.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.