Courseiva

Databricks-ML-Assoc Databricks Machine Learning Practice Question

A data engineer is working with a large Delta table and wants to optimize it for machine learning feature engineering. They frequently filter data by a column named 'event_date' and join on a column named 'user_id'. Which Delta Lake feature should they use to improve query performance?

⚠ Common exam trap

The trap here is assuming that partitioning always improves performance, but for high-cardinality columns it can create too many small files and worsen performance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Z-ORDER BY on 'user_id' and 'event_date'.

Z-ORDER BY clusters data on the specified columns, enabling efficient data skipping for filters and joins. Since the engineer filters by 'event_date' and joins on 'user_id', Z-ORDER on both columns will significantly improve performance. Partitioning on high-cardinality 'user_id' is inefficient, and CDF or simple compaction do not address the need for co-locating related data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enabling change data feed (CDF).

    Why it's wrong here

    Change data feed tracks row-level changes for downstream consumption, not for improving query performance on static data. It adds overhead and is unrelated to optimizing reads for feature engineering. CDF is useful for incremental processing but does not enhance data skipping or clustering for analytical queries.

  • ✗

    Using OPTIMIZE with file compaction.

    Why it's wrong here

    OPTIMIZE compacts small files into larger ones, which can improve read performance, but without Z-ORDER it does not co-locate data by the filter columns. It may reduce the number of files but does not provide data skipping based on 'event_date' or 'user_id'. For the described access pattern, Z-ORDER is more effective.

  • ✗

    Partitioning by 'user_id'.

    Why it's wrong here

    Partitioning by 'user_id' could create a huge number of small partitions if user_id has high cardinality, leading to performance degradation and metadata overhead. It also does not help with filtering by 'event_date'. While partitioning can be beneficial for low-cardinality columns, it is not ideal for high-cardinality columns like user_id, and it does not address the date filtering need.

  • ✓

    Z-ORDER BY on 'user_id' and 'event_date'.

    Why this is correct

    Z-ORDER BY co-locates related data in the same set of files, improving data skipping for queries that filter on the specified columns. By clustering on 'user_id' and 'event_date', queries that filter by 'event_date' and join on 'user_id' can skip irrelevant files, reducing I/O and speeding up feature engineering. This is the recommended optimization for such access patterns.

About these practice questions

One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.