Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

A data engineer is building a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The pipeline must tolerate late-arriving data up to 10 minutes and update aggregations accordingly. The engineer wants to use a watermark on the event-time column. Which code snippet correctly applies the watermark and performs a 5-minute tumbling window aggregation?

⚠ Common exam trap

The trap here is assuming the watermark can be applied after the aggregation or that its duration must match the window size.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.withWatermark("event_time", "10 minutes").groupBy(window("event_time", "5 minutes")).count()

The correct approach is to apply the watermark on the event-time column before the aggregation, with a delay that matches the late-data tolerance (10 minutes). The aggregation must use a window of the desired size (5 minutes). Placing the watermark after the aggregation or using mismatched durations leads to incorrect results or errors.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    df.withWatermark("event_time", "10 minutes").groupBy("event_time").count()

    Why it's wrong here

    This applies a watermark but groups by the raw event_time column instead of a window. Without a window, the aggregation would produce one row per distinct timestamp, not a 5-minute tumbling window. This fails to meet the windowing requirement and would result in a non-windowed aggregation.

  • ✓

    df.withWatermark("event_time", "10 minutes").groupBy(window("event_time", "5 minutes")).count()

    Why this is correct

    This correctly applies a watermark of 10 minutes on the event_time column, allowing late data up to that threshold, and then groups by a 5-minute tumbling window. The watermark is set before the aggregation, which is required for streaming aggregations. This matches the requirement to tolerate late data and compute windowed counts.

  • ✗

    df.withWatermark("event_time", "5 minutes").groupBy(window("event_time", "10 minutes")).count()

    Why it's wrong here

    The watermark is set to 5 minutes, but the requirement is to tolerate late data up to 10 minutes. Additionally, the window duration is 10 minutes instead of the required 5 minutes. This mismatch would drop data arriving between 5 and 10 minutes late and produce incorrect window sizes.

  • ✗

    df.groupBy(window("event_time", "5 minutes")).count().withWatermark("event_time", "10 minutes")

    Why it's wrong here

    The watermark is applied after the aggregation, which is invalid in Structured Streaming. Watermarks must be defined on the input DataFrame before any stateful operation like groupBy. This ordering would cause an AnalysisException because the watermark cannot be applied to the result of an aggregation in this manner.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.