Databricks-Spark-Assoc Structured Streaming Practice Question
A developer is building a Structured Streaming job in PySpark that reads from a Kafka topic and writes to a Delta Lake table. The job uses `outputMode("append")` and a 10-minute watermark on the event-time column `event_time`. A batch of late data arrives with events whose `event_time` is older than the watermark. What happens to these late events?
⚠ Common exam trap
The trap here is assuming that late data is either reprocessed or causes an error, rather than being silently dropped as the watermark dictates.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The late events are dropped and not included in the output.
When a watermark is set on an event-time column, Structured Streaming uses it to determine when data is too late to be included in an aggregation. Events with timestamps older than the watermark are dropped. This behavior is intentional to bound state and ensure timely results. The other options describe mechanisms that are not part of the watermarking feature, such as buffering, erroring, or automatic dead-lettering.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The late events are buffered and processed once the watermark advances.
Why it's wrong here
Structured Streaming does not buffer late events for future processing. Once the watermark passes the event time of a record, that record is considered too late and is dropped, not held. Buffering would defeat the purpose of watermarking, which is to limit state growth. The engine uses the watermark to decide when to finalize aggregations and discard old state, so late events cannot be processed later. This option mischaracterizes the watermark mechanism.
- ✓
The late events are dropped and not included in the output.
Why this is correct
In Structured Streaming, when a watermark is defined on an event-time column, any event whose event-time timestamp is older than the current watermark (max event time seen minus the watermark delay) is considered too late and is dropped from the aggregation. This is the documented behavior: watermarking allows the engine to discard late data to bound state. Since the job uses append mode with a watermark, late events beyond the threshold are excluded from the result set, which is the correct outcome here.
- ✗
The late events cause the query to fail with a watermark violation error.
Why it's wrong here
Late events do not cause a failure; they are simply dropped. Structured Streaming is designed to handle out-of-order and late data gracefully when a watermark is defined. A watermark violation error is not a standard Spark exception. The query continues running, and late events are ignored. This option invents an error that does not exist, confusing watermarking with a strict validation rule.
- ✗
The late events are written to a separate dead-letter queue automatically.
Why it's wrong here
Structured Streaming does not automatically route late events to a dead-letter queue. There is no built-in dead-letter mechanism for late data. If you need to capture late events, you must implement custom logic, such as using `foreachBatch` to write them elsewhere, but that is not automatic. This option assumes a feature that does not exist in the Spark Structured Streaming API, leading to an incorrect expectation.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.