Courseiva
Data Preparation →hardMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

A team is preparing a Delta table of product descriptions for a RAG application. The table receives continuous upserts from a streaming pipeline, and the embedding job reads the table every hour. Engineers notice the embedding job reprocesses every row on each run even though only a few rows change. Which change should be made to the source table to let the embedding job process only new or updated rows?

⚠ Common exam trap

Candidates often confuse read-performance optimizations such as partitioning or OPTIMIZE with change tracking, when only Change Data Feed exposes which rows were inserted, updated, or deleted.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable Change Data Feed on the Delta table and have the embedding job read the change feed since the last processed version.

Change Data Feed captures row-level changes with commit versions, allowing a downstream job to read only the inserts, updates, and deletes that occurred since its last checkpoint. This directly solves the problem of reprocessing unchanged rows. Partitioning, compaction, and timestamp filtering improve read characteristics but do not expose which rows changed, so they cannot deliver incremental processing.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Partition the Delta table by the product category column so the embedding job reads fewer files.

    Why it's wrong here

    Partitioning by category reduces scan scope only when queries filter on that column, but the embedding job still has no way to distinguish changed rows from unchanged ones within a partition. It cannot avoid reprocessing rows that were already embedded, so the core problem remains.

  • ✓

    Enable Change Data Feed on the Delta table and have the embedding job read the change feed since the last processed version.

    Why this is correct

    Change Data Feed records row-level inserts, updates, and deletes with commit versions, so the embedding job can read only changes since its last checkpoint. This avoids reprocessing unchanged rows on every run. It is the intended mechanism for incremental downstream consumption of a Delta table that receives continuous upserts.

  • ✗

    Convert the table to a view that filters rows by the current timestamp so only recent rows appear.

    Why it's wrong here

    Filtering by timestamp assumes the table has a reliable modification timestamp column, which an upsert pipeline may not maintain, and it would miss updates to older rows. It also cannot distinguish which rows were already embedded, so it does not provide reliable incremental processing.

  • ✗

    Run OPTIMIZE with Z-ORDER on the product identifier column to compact small files before each embedding run.

    Why it's wrong here

    Compaction and Z-ORDER improve read performance by reducing file count and clustering related data, but they do not expose which rows changed. The embedding job would still read the full table each hour, so reprocessing is not eliminated.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.