Courseiva
Data Preparation →mediumMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

A Generative AI engineer is building a Delta Live Tables pipeline that ingests raw JSON event logs into a bronze table, then uses ai_query to classify each event's free-text field. The classification call is expensive, so the engineer wants to avoid re-running it on events that have already been processed in previous pipeline updates. The source table is append-only and new events arrive continuously. Which Delta Live Tables feature should the engineer configure on the bronze table to prevent reprocessing of previously ingested rows?

⚠ Common exam trap

The trap here is assuming that a table property or filter clause can prevent reprocessing, when incremental read semantics in Delta Live Tables come from declaring the dataset as a streaming table.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the bronze table as a streaming table so it processes only new data on each pipeline update.

An append-only bronze ingestion table should be defined as a streaming table so each pipeline update consumes only new source records. This makes expensive operations like ai_query run once per new event rather than on the full history. Materialized views and batch-style definitions re-read all source data, and storage properties like Auto Optimize only affect file layout, not incremental read behavior.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure the bronze table as a streaming table so it processes only new data on each pipeline update.

    Why this is correct

    Streaming tables in Delta Live Tables read incrementally, so on each pipeline update only newly arrived source rows flow through the ai_query classification. Previously ingested events are not re-read, which is exactly what avoids paying for repeated classification calls. This is the standard pattern for append-only ingestion with expensive downstream transformations.

  • ✗

    Add a QUALIFY clause that filters out rows where the classification column is already populated.

    Why it's wrong here

    QUALIFY filters on window function results within a single query, not across pipeline runs. It has no memory of what was processed in a prior update, so it cannot skip already-classified events. The engine would still read every source row and re-evaluate the ai_query expression, defeating the cost-avoidance goal.

  • ✗

    Enable change data capture by setting pipelines.cdcEnabled to true on the bronze table.

    Why it's wrong here

    There is no pipelines.cdcEnabled table property; change data capture in DLT is expressed through the APPLY CHANGES APIs, not a boolean flag. Even if CDC were configured, it governs how updates and deletes are applied to a target table, not how a bronze ingestion table avoids re-reading unchanged append-only source rows.

  • ✗

    Set the table property pipelines.autoOptimize.managed to true on the bronze table.

    Why it's wrong here

    Auto Optimize manages file compaction and layout for Delta tables; it does not change which source rows a streaming DLT table reads on subsequent updates. It cannot prevent ai_query from executing again on old events because the same input rows are still read from the source. It is a storage-optimization property, not an incremental-read control.

About these practice questions

This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.