Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

A Data Engineer is using Delta Live Tables to process a stream of financial transactions. The pipeline must ensure that each `transaction_id` appears only once in the target table, even if the source stream contains duplicates due to at-least-once ingestion. The engineer wants to use the `APPLY CHANGES` API. Which combination of settings will achieve this with minimal data loss?

⚠ Common exam trap

The trap here is assuming that SCD Type 2 is needed for deduplication, when in fact it preserves history and creates multiple rows per key.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use `APPLY CHANGES INTO target` with `KEYS (transaction_id)` and `SEQUENCE BY timestamp`, and set `STORED AS SCD TYPE 1`.

To deduplicate and keep only the latest record per `transaction_id`, `APPLY CHANGES` with SCD Type 1 and a sequence column is correct. SCD Type 1 performs upserts, overwriting existing rows, so each key appears once. SCD Type 2 would keep history, resulting in multiple rows. The sequence column ensures the latest record is applied deterministically.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use `APPLY CHANGES INTO target` with `KEYS (transaction_id)` and `SEQUENCE BY timestamp`, and set `STORED AS SCD TYPE 1`.

    Why this is correct

    SCD Type 1 with `APPLY CHANGES` uses the specified keys to upsert records, keeping only the latest version based on the sequence column. This ensures each `transaction_id` appears once, with the most recent data. It handles duplicates by overwriting existing rows, which is appropriate for deduplication when the latest record is desired.

  • ✗

    Use `APPLY CHANGES INTO target` with `KEYS (transaction_id)` and omit the `SEQUENCE BY` clause, and set `STORED AS SCD TYPE 1`.

    Why it's wrong here

    Without `SEQUENCE BY`, Delta Live Tables cannot determine the order of updates for the same key, which can lead to non-deterministic results when duplicates arrive out of order. The sequence column is essential to correctly identify the latest record. Omitting it may cause incorrect data to be retained, failing the deduplication goal.

  • ✗

    Use `APPLY CHANGES INTO target` with `KEYS (transaction_id)` and `SEQUENCE BY timestamp`, and set `STORED AS SCD TYPE 2`.

    Why it's wrong here

    SCD Type 2 retains history by creating new rows for changes, so the target table would contain multiple rows per `transaction_id` over time. This violates the requirement that each ID appears only once. SCD Type 2 is for tracking historical changes, not for deduplication to a single current record.

  • ✗

    Use `APPLY CHANGES INTO target` with `KEYS (transaction_id)` and `SEQUENCE BY timestamp`, and set `STORED AS SCD TYPE 2` with `IGNORE NULLS`.

    Why it's wrong here

    SCD Type 2 still creates multiple rows per key to track history, even with `IGNORE NULLS`. The requirement is to have each `transaction_id` appear only once, so SCD Type 2 is fundamentally unsuitable. `IGNORE NULLS` only affects how nulls are handled in the sequence column, not the duplication issue.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.