Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

A data engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data and update already processed aggregates. Which watermark strategy should be used to allow updates to aggregates while bounding state store growth?

⚠ Common exam trap

The trap here is assuming that any watermark will automatically handle late data, but the delay must be set based on the actual lateness characteristics; too short a delay drops data, too long a delay grows state unnecessarily.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a watermark on the event-time column with a delay that accommodates the maximum expected lateness, and set the output mode to 'update'.

The correct answer is to use an event-time watermark with a delay that matches the maximum expected lateness and to use the 'update' output mode. This allows late data to be incorporated into aggregates while bounding state store growth. The watermark tells the engine how long to wait for late data before finalizing a window, and 'update' mode emits incremental updates, which is ideal for updating aggregates. Other options either use inappropriate time semantics or set delays that are too short.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a watermark based on processing time rather than event time, with a 10-minute delay.

    Why it's wrong here

    Watermarks based on processing time do not handle out-of-order or late-arriving data effectively because they rely on the system clock, not the event timestamps. Late data would be dropped or processed incorrectly. Structured Streaming watermarks are designed to work with event-time columns to track progress and allow late data within a threshold. Using processing time would not correctly bound state based on event-time lateness and could lead to incorrect aggregates.

  • ✗

    Use a watermark on the event-time column with a delay of 0 seconds to ensure all data is processed immediately.

    Why it's wrong here

    A watermark delay of 0 seconds means that any data arriving after the current maximum event time is considered too late and will be dropped. This would cause almost all late-arriving data to be discarded, leading to incomplete aggregates. While it would bound state growth aggressively, it sacrifices correctness for late data. The scenario requires handling late-arriving data, so a zero delay is inappropriate.

  • ✓

    Use a watermark on the event-time column with a delay that accommodates the maximum expected lateness, and set the output mode to 'update'.

    Why this is correct

    This approach correctly uses event-time watermarks to bound state by allowing late data up to a specified delay. The 'update' output mode emits only updated aggregates, which is efficient for updating results. By setting the delay to the maximum expected lateness, the engineer ensures that late data within that window is processed, and state older than the watermark is dropped. This balances correctness and resource usage in a streaming aggregation.

  • ✗

    Use a static watermark of 1 hour by setting the watermark delay to '1 hour' on the event-time column.

    Why it's wrong here

    A static watermark of 1 hour would drop any events that arrive later than 1 hour after the maximum event time seen so far. In a scenario with late-arriving data that may exceed 1 hour, this would incorrectly discard valid updates. Watermarks are meant to bound state, but a fixed 1-hour delay may be too short or too long; it does not adapt to actual lateness patterns. The engineer needs a watermark that balances state size and data completeness, not an arbitrary fixed value.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.