You are implementing a Spark Structured Streaming job in Azure Databricks that reads from an Azure Event Hubs topic. The job must handle late-arriving data up to 10 minutes and produce aggregated results every 5 minutes. You need to configure the watermark and window. Which code snippet should you use?
This snippet sets a watermark of 10 minutes to allow late data up to that threshold and groups by a 5-minute tumbling window. The watermark defines how long the system waits for late events before finalizing a window, and the window defines the aggregation interval. This matches the requirement to handle late data up to 10 minutes and produce results every 5 minutes, assuming output mode is set appropriately.
Why this answer
The correct configuration requires a watermark of 10 minutes to accommodate late data up to that limit, and a 5-minute tumbling window to produce results every 5 minutes. The withWatermark method sets the watermark on the event time column, and the window function with a single duration creates a tumbling window. The output mode must be set to update or append depending on the sink requirements.
Exam trap
The trap here is swapping the watermark duration and window duration, or using a sliding window when a tumbling window is needed.