Databricks-DE-Pro Data Modelling Practice Question
A financial institution uses a Databricks Lakehouse with a Silver table transactions that is partitioned by transaction_date. The table is frequently queried with filters on transaction_date and account_id. The data engineering team notices that queries filtering on account_id are slow because they scan all partitions. They want to optimize the table to accelerate these queries without repartitioning. Which Delta Lake feature should they use?
⚠ Common exam trap
The trap here is thinking that adding another partition column will solve the problem, but high-cardinality columns are better handled with Z-ORDER to avoid the small file problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Z-ORDER BY account_id
Z-ORDER BY account_id clusters data within each partition by account_id, allowing Delta Lake to skip files that do not contain the filtered account_id values. This accelerates queries filtering on account_id without altering the existing partitioning by transaction_date. Other options either do not improve performance or introduce negative side effects.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Partition the table by account_id as well.
Why it's wrong here
Adding account_id as a partition column would create a multi-level partitioning scheme, but account_id likely has high cardinality, leading to many small partitions and performance degradation. It also requires rewriting the table, which the team wants to avoid. Partitioning by both columns is not recommended.
- ✗
Convert the table to a Parquet table and use predicate pushdown.
Why it's wrong here
Parquet supports predicate pushdown, but converting from Delta Lake would lose ACID transactions, time travel, and other benefits. The existing partitioning by transaction_date already provides predicate pushdown on that column. This change would not specifically accelerate filters on account_id and introduces drawbacks.
- ✓
Z-ORDER BY account_id
Why this is correct
Z-ORDER BY account_id will co-locate similar account_id values within each partition, enabling data skipping for filters on account_id. Since the table is already partitioned by transaction_date, this adds efficient skipping for account_id without changing the partitioning scheme. This directly addresses the slow queries filtering on account_id.
- ✗
Enable change data feed on the table.
Why it's wrong here
Change data feed tracks row-level changes for downstream consumers, but it does not improve query performance for filtering on account_id. It adds overhead and is unrelated to data skipping or clustering. Thus, it does not solve the slow query problem.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.