Courseiva

Databricks-DE-Pro Data Ingestion and Acquisition Practice Question

A data engineer is building a streaming ingestion pipeline from Apache Kafka to a Delta table. The pipeline must perform deduplication on a unique event_id field and handle late-arriving data. The engineer wants to use Structured Streaming with a watermark of 10 minutes. Which of the following approaches correctly implements deduplication and watermarking?

⚠ Common exam trap

The trap here is assuming that dropDuplicates alone can handle late data without a watermark, but without a watermark, state grows indefinitely and late data may be dropped incorrectly.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use dropDuplicates with event_id and set the watermark on the event timestamp column.

The correct approach is to use dropDuplicates on the unique event_id and set a watermark on the event timestamp column. This combination allows the stream to deduplicate events within the watermark window, effectively handling late-arriving data while bounding state. It is the recommended pattern for streaming deduplication in Databricks Structured Streaming.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use dropDuplicates with event_id and the event timestamp column without setting a watermark.

    Why it's wrong here

    Without a watermark, dropDuplicates on event_id and timestamp will only deduplicate within the same micro-batch or across the entire stream if state is maintained indefinitely, which can lead to unbounded state growth. The watermark is essential to bound the state and handle late data. This approach does not correctly handle late-arriving data and may cause memory issues.

  • ✗

    Use a windowed aggregation on event_id and timestamp, then filter out duplicates using row_number.

    Why it's wrong here

    Using windowed aggregation and row_number can deduplicate, but it is more complex and may not be as efficient as dropDuplicates. It also requires careful handling of late data and state management. The dropDuplicates with watermark is simpler and directly addresses the requirement.

  • ✗

    Use foreachBatch to manually deduplicate by querying the Delta table for existing event_id values.

    Why it's wrong here

    While foreachBatch can be used for custom deduplication, it is not the recommended approach for streaming deduplication because it requires querying the Delta table for each batch, which can be inefficient and may not handle late data correctly. The built-in dropDuplicates with watermark is more efficient and designed for this purpose.

  • ✓

    Use dropDuplicates with event_id and set the watermark on the event timestamp column.

    Why this is correct

    This approach correctly uses dropDuplicates on the unique event_id to remove duplicates within the watermark window. The watermark on the event timestamp column allows the stream to handle late data by tracking how late data can arrive. This is the standard pattern for deduplication in Structured Streaming with watermarking, ensuring exactly-once semantics when combined with Delta Lake.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.