Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

A Data Engineer is using Delta Live Tables to process a stream of user events. The `user_id` column should be unique in the target table, but the source may contain duplicate events due to retries. The engineer wants to keep only the latest event for each `user_id` based on the `event_timestamp`. Which Delta Live Tables feature should be used?

⚠ Common exam trap

The trap here is attempting to use window functions or dropDuplicates for streaming deduplication, which are either unsupported or limited in streaming contexts.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use `APPLY CHANGES INTO target` with `KEYS (user_id)` and `SEQUENCE BY event_timestamp`, and set `STORED AS SCD TYPE 1`.

To deduplicate and keep only the latest event per `user_id`, `APPLY CHANGES` with SCD Type 1 and a sequence column is the correct Delta Live Tables feature. It performs upserts based on the key and sequence, ensuring one row per key with the latest data. SCD Type 2 would keep history, and window functions or dropDuplicates are not suitable for streaming deduplication.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use `APPLY CHANGES INTO target` with `KEYS (user_id)` and `SEQUENCE BY event_timestamp`, and set `STORED AS SCD TYPE 2`.

    Why it's wrong here

    SCD Type 2 would retain all historical events for each `user_id`, resulting in multiple rows per user. The requirement is to have only the latest event, so SCD Type 2 is not suitable. It is designed for tracking history, not for deduplication to a single current record.

  • ✗

    Use `@dp.expect_or_drop("unique_user", "user_id IS NOT NULL")` and then apply `dropDuplicates(["user_id"])` on the resulting DataFrame.

    Why it's wrong here

    `dropDuplicates` on a streaming DataFrame without a watermark is not supported for streaming deduplication; it requires a watermark and will only deduplicate within the watermark window. It does not guarantee global uniqueness and may not keep the latest event. This approach does not reliably achieve the goal.

  • ✓

    Use `APPLY CHANGES INTO target` with `KEYS (user_id)` and `SEQUENCE BY event_timestamp`, and set `STORED AS SCD TYPE 1`.

    Why this is correct

    This configuration uses `APPLY CHANGES` to upsert records by `user_id`, keeping the latest based on `event_timestamp`. SCD Type 1 ensures only the current record is stored, so each `user_id` appears once. It handles duplicates by overwriting with the latest event, which is exactly what is needed.

  • ✗

    Use `@dp.expect_or_drop("unique_user", "row_number() OVER (PARTITION BY user_id ORDER BY event_timestamp DESC) = 1")` in a streaming table definition.

    Why it's wrong here

    Window functions like `row_number()` are not supported in streaming table expectations because they require a shuffle and are not allowed in streaming queries. This approach would fail. Deduplication in streaming typically requires `APPLY CHANGES` or `dropDuplicates` with watermarking, but not window functions in expectations.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.