Courseiva

Databricks-DE-Pro Data Ingestion and Acquisition Practice Question

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the messages can arrive out of order by up to 10 minutes. The engineer wants to perform time-windowed aggregations on the ingested data while minimizing state store overhead. Which approach should be used to handle the out-of-order data correctly?

⚠ Common exam trap

It's easy for candidates to confuse Kafka consumer properties or offset management with event-time processing, leading to solutions that do not actually handle late-arriving data or manage state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use withWatermark on the event-time column extracted from the Kafka value and set the watermark to 10 minutes.

To handle out-of-order data in Structured Streaming, you must use event-time processing with watermarks. Extracting the event-time column from the Kafka value and applying a watermark of 10 minutes allows the engine to wait for late data up to that threshold and then clean up state. This minimizes state store overhead and ensures correct time-windowed aggregations. Other options do not address event-time semantics or state management.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use withWatermark on the event-time column extracted from the Kafka value and set the watermark to 10 minutes.

    Why this is correct

    Watermarking allows the stream to handle late data by defining a threshold beyond which late events are dropped. By extracting the event-time column from the Kafka value and applying a 10-minute watermark, the aggregation can correctly include events that arrive within that window. This minimizes state store overhead because the watermark instructs the engine to clean up old state. It is the standard approach for out-of-order data in Structured Streaming.

  • ✗

    Configure the Kafka source with 'startingOffsets' set to 'latest' and enable auto-commit of offsets.

    Why it's wrong here

    Starting offsets and auto-commit control where the stream begins reading and whether offsets are committed automatically. They do not influence event-time processing or out-of-order handling. Auto-commit can actually lead to data loss if not managed properly. This option is about offset management, not about correctly aggregating time-windowed data with late arrivals. It does not solve the problem.

  • ✗

    Repartition the stream by the Kafka partition ID and sort within each partition using a custom foreachBatch function.

    Why it's wrong here

    Repartitioning and sorting within foreachBatch might reorder data within a micro-batch, but it does not provide a mechanism for handling late events across batches. The state store would still grow unbounded because there is no watermark to expire old state. This approach is complex and does not leverage Structured Streaming's built-in event-time processing. It fails to address the core requirement of handling out-of-order data with minimal state overhead.

  • ✗

    Set the Kafka consumer property 'isolation.level' to 'read_committed' to ensure only committed messages are processed.

    Why it's wrong here

    The isolation.level property controls whether the consumer reads only committed messages in a transactional Kafka producer scenario. It does not address out-of-order arrival or event-time processing. While it can prevent reading uncommitted data, it has no effect on watermarking or late data handling. This option is unrelated to the problem of ordering and would not help with time-windowed aggregations.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.