Courseiva

Databricks-DE-Assoc Data Transformation and Modeling Practice Question

A data engineer is working with a Delta table that contains a column 'timestamp' of type timestamp. The table is partitioned by date. The engineer needs to run a query that filters on a specific date range and also on a high-cardinality column 'user_id'. The query is performing poorly. Which optimization technique should the engineer apply to improve query performance?

⚠ Common exam trap

The trap here is assuming that caching or repartitioning will solve selective query performance, when the real issue is data layout for skipping on a high-cardinality column.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Z-ORDER BY on the 'user_id' column when writing the Delta table.

Z-ORDER BY is designed to optimize data skipping for high-cardinality columns by clustering data with similar values in the same files. This reduces the number of files scanned when filtering on that column. Partitioning on date already helps with date-range filters, but Z-Ordering on user_id addresses the high-cardinality filter. Caching, repartitioning, or converting to Parquet do not provide the same targeted benefit for this scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Z-ORDER BY on the 'user_id' column when writing the Delta table.

    Why this is correct

    Z-ORDER BY is a technique that colocates related data in the same set of files, improving data skipping for queries that filter on the Z-ordered columns. By Z-Ordering on 'user_id', the engineer can reduce the amount of data scanned when filtering on that high-cardinality column. This is especially effective when combined with partitioning on date, as it further optimizes within each partition. It is a recommended best practice for Delta Lake performance tuning.

  • ✗

    Convert the Delta table to Parquet format to leverage predicate pushdown.

    Why it's wrong here

    Delta Lake already supports predicate pushdown and data skipping, often more efficiently than plain Parquet due to its transaction log and statistics. Converting to Parquet would lose Delta's advanced features like ACID transactions, time travel, and Z-Ordering. Predicate pushdown is not exclusive to Parquet; Delta Lake also provides it. Thus, this change would not improve performance and would sacrifice important capabilities.

  • ✗

    Enable Delta Lake caching by running CACHE TABLE on the Delta table.

    Why it's wrong here

    Caching the entire Delta table in memory can improve performance for repeated queries, but it is not a targeted optimization for filtering on high-cardinality columns. It consumes significant memory and may not be feasible for large tables. Moreover, caching does not help with data skipping or predicate pushdown, which are more effective for selective queries. This approach is typically used for frequently accessed tables, not for optimizing a specific query pattern.

  • ✗

    Repartition the Delta table by the 'timestamp' column.

    Why it's wrong here

    Repartitioning by 'timestamp' would reorganize the data but is unlikely to improve filtering on 'user_id'. Repartitioning is a physical layout change that can help with parallelism but does not provide data skipping benefits for high-cardinality columns. Since the table is already partitioned by date, further repartitioning by timestamp may lead to many small files and degrade performance. This is not the appropriate optimization for the given query pattern.

About these practice questions

One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.