Courseiva
Data Modelling →mediumMultiple Choice

Databricks-DE-Pro Data Modelling Practice Question

A Databricks workspace has a Delta table 'transactions' partitioned by 'txn_date'. Analysts frequently run queries that filter on 'txn_date' but also occasionally filter on 'account_id' alone. The table has 10 TB of data, and the team wants to improve performance for the 'account_id' queries without changing the partitioning scheme. Which Delta feature should they implement?

⚠ Common exam trap

The trap here is assuming that adding another partition column will automatically improve filtering, when high-cardinality partitioning often causes small-file problems and slower queries.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Z-ORDER BY account_id

Z-Ordering is the correct approach because it reorganizes data within existing partitions to improve data skipping for the specified columns without altering the partition structure. It is specifically designed to optimize queries on high-cardinality columns like account_id. Partitioning by account_id would cause scalability issues, CDF is for change tracking, and Hive tables lack Delta optimizations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Z-ORDER BY account_id

    Why this is correct

    Z-Ordering colocates related data in the same set of files, so queries filtering on account_id can skip many files. It is ideal when the table is already partitioned by txn_date and you need to optimize a secondary column without repartitioning. Running OPTIMIZE ... Z-ORDER BY account_id will reorganize data within each partition, improving data skipping for account_id filters.

  • ✗

    Partition the table by account_id as well as txn_date

    Why it's wrong here

    Adding account_id as a partition column would create a huge number of small partitions (high cardinality), leading to many tiny files and metadata overhead. This would degrade performance for all queries, not improve account_id filters. Partitioning is best for low-cardinality columns, and account_id likely has millions of distinct values.

  • ✗

    Enable Delta Lake change data feed

    Why it's wrong here

    Change data feed (CDF) tracks row-level changes for downstream consumption, not for accelerating query performance. It adds metadata and does not reorganize data for skipping. Enabling CDF would not help queries filtering on account_id; it is designed for CDC scenarios, not for optimizing existing read patterns.

  • ✗

    Convert the table to a Hive table

    Why it's wrong here

    Converting to a Hive table would lose Delta Lake's advanced features like ACID transactions, time travel, and data skipping. Hive tables do not support Z-Ordering or Delta-specific optimizations. This would likely worsen performance and is not a recommended approach for improving query speed in Databricks.

About these practice questions

One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.