Databricks-DE-Pro Data Modelling Practice Question
A Databricks workspace has a Delta table 'transactions' partitioned by 'txn_date'. Analysts frequently run queries that filter on 'txn_date' but also occasionally filter on 'account_id' alone. The table has 10 TB of data, and the team wants to improve performance for the 'account_id' queries without changing the partitioning scheme. Which Delta feature should they implement?
⚠ Common exam trap
The trap here is assuming that adding another partition column will automatically improve filtering, when high-cardinality partitioning often causes small-file problems and slower queries.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Z-ORDER BY account_id
Z-Ordering is the correct approach because it reorganizes data within existing partitions to improve data skipping for the specified columns without altering the partition structure. It is specifically designed to optimize queries on high-cardinality columns like account_id. Partitioning by account_id would cause scalability issues, CDF is for change tracking, and Hive tables lack Delta optimizations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Z-ORDER BY account_id
Why this is correct
Z-Ordering colocates related data in the same set of files, so queries filtering on account_id can skip many files. It is ideal when the table is already partitioned by txn_date and you need to optimize a secondary column without repartitioning. Running OPTIMIZE ... Z-ORDER BY account_id will reorganize data within each partition, improving data skipping for account_id filters.
- ✗
Partition the table by account_id as well as txn_date
Why it's wrong here
Adding account_id as a partition column would create a huge number of small partitions (high cardinality), leading to many tiny files and metadata overhead. This would degrade performance for all queries, not improve account_id filters. Partitioning is best for low-cardinality columns, and account_id likely has millions of distinct values.
- ✗
Enable Delta Lake change data feed
Why it's wrong here
Change data feed (CDF) tracks row-level changes for downstream consumption, not for accelerating query performance. It adds metadata and does not reorganize data for skipping. Enabling CDF would not help queries filtering on account_id; it is designed for CDC scenarios, not for optimizing existing read patterns.
- ✗
Convert the table to a Hive table
Why it's wrong here
Converting to a Hive table would lose Delta Lake's advanced features like ACID transactions, time travel, and data skipping. Hive tables do not support Z-Ordering or Delta-specific optimizations. This would likely worsen performance and is not a recommended approach for improving query speed in Databricks.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.