Databricks-DE-Pro Cost and Performance Optimization Practice Question
A data engineering team is running a nightly batch job on a Databricks job cluster that processes a 10 TB Delta table. The job reads the entire table, performs transformations, and writes results to another Delta table. The team notices that the job takes 4 hours and consumes significant DBUs. They want to reduce runtime and cost without changing the business logic. The table is partitioned by ingestion date, but queries often filter on a high-cardinality column 'customer_id'. Which optimization technique is most appropriate to improve performance and reduce cost?
⚠ Common exam trap
The trap here is assuming that adding more cluster resources or caching will solve performance issues without addressing data layout, which is often the root cause.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Optimize file layout using Z-ORDER BY on the 'customer_id' column.
Z-ORDER BY on 'customer_id' colocates related data, enabling Delta Lake's data skipping to prune files during reads. This reduces I/O and compute, directly lowering runtime and DBU cost. Other options either increase cost, are impractical, or degrade performance. The optimization is persistent and does not require changes to business logic, making it ideal for this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable Delta Lake caching by running CACHE TABLE on the source table before the job.
Why it's wrong here
Caching the entire 10 TB table in memory is impractical because it would require a cluster large enough to hold the data, increasing cost. Also, caching is ephemeral and not persisted across job runs, so the nightly job would still need to read from storage. This does not address the underlying data layout issue and may cause memory pressure, leading to spilling and slower performance.
- ✗
Increase the cluster size by adding more worker nodes to the job cluster.
Why it's wrong here
Adding more workers may speed up the job but also increases DBU consumption, raising cost. Without addressing data layout, the job still reads unnecessary data, so scaling out is inefficient. This approach does not solve the root cause of slow performance due to poor data skipping on the high-cardinality filter column, and it may not yield linear speedup due to I/O bottlenecks.
- ✓
Optimize file layout using Z-ORDER BY on the 'customer_id' column.
Why this is correct
Z-ORDER BY clusters data on the 'customer_id' column, improving data skipping for queries that filter on that column. This reduces the amount of data read during the nightly job, leading to faster runtime and lower DBU consumption. It is a cost-effective optimization because it reorganizes existing data without changing business logic, and the benefits persist across runs until the data is rewritten.
- ✗
Convert the table to a Parquet table and use partition pruning on 'customer_id'.
Why it's wrong here
Converting to Parquet loses Delta Lake features like ACID transactions and time travel, which may be required. Partitioning by 'customer_id' would create too many small partitions due to high cardinality, causing performance degradation. This approach is not cost-effective and may break downstream processes. It also does not leverage Databricks optimizations like Z-ORDER or data skipping.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.