Courseiva
Structured Streaming →mediumMultiple Choice

Databricks-Spark-Assoc Structured Streaming Practice Question

You are processing a streaming dataset of sensor readings. You need to calculate the average temperature every 10 minutes, allowing data to arrive up to 2 minutes late. Which windowing approach correctly handles this requirement in Structured Streaming?

⚠ Common exam trap

Candidates often forget to pair the window function with a watermark, which causes Spark to throw an analysis error or accumulate unbounded state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use window(timestamp, '10 minutes') with withWatermark('timestamp', '2 minutes').

Event-time windowing with watermarks is essential for handling late-arriving data in stream processing. By defining a window duration of 10 minutes and a watermark delay of 2 minutes, Spark maintains state for late events while discarding data older than the watermark threshold. This ensures the output remains accurate even when network latency or ingestion bottlenecks occur, preventing unbounded state growth in the processing engine.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use window(timestamp, '10 minutes') without a watermark.

    Why it's wrong here

    Without a watermark, the streaming query will retain all historical state indefinitely to account for potentially late data. This leads to a continuous increase in memory usage, eventually causing an OutOfMemoryError as the state store grows beyond the available cluster resources, making it unsuitable for long-running production streaming applications.

  • ✗

    Use window(timestamp, '10 minutes') with withWatermark('timestamp', '10 minutes').

    Why it's wrong here

    Setting the watermark delay to 10 minutes is excessively high for a 2-minute late-data requirement. This configuration forces the engine to retain state for much longer than necessary, significantly increasing memory pressure on the executor nodes and introducing unnecessary latency in the final calculation results for the windowed output.

  • ✓

    Use window(timestamp, '10 minutes') with withWatermark('timestamp', '2 minutes').

    Why this is correct

    This configuration perfectly matches the requirement by grouping events into 10-minute buckets while allowing a 2-minute buffer for late data. The watermark correctly signals to Spark that state for windows older than 2 minutes from the maximum event time can be safely cleared, optimizing memory usage and ensuring production stability.

  • ✗

    Use a trigger interval of 2 minutes with no watermark.

    Why it's wrong here

    Trigger intervals control how often the streaming query executes, not how late data is processed. Without a watermark, the state management remains unbounded, and the query will not inherently handle late data correctly based on event time. This approach fails to provide the state-clearing mechanism required for long-running streaming jobs.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.