Databricks-Spark-Assoc Structured Streaming Practice Question
An engineer is developing a Structured Streaming job that reads from an Apache Kafka source and writes the output continuously toDelta Lake using outputMode("append"). The stream occasionally experiences late-arriving data. Which downstream behavior can the engineer expect regarding the Delta Lake table?
⚠ Common exam trap
Candidates often assume that late-arriving data is automatically updated in the destination table or buffered indefinitely, forgetting that watermarks strictly drop data that falls behind the threshold in append mode.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Late-arriving data that falls behind the specified watermark threshold is dropped and never written to the Delta table.
Streaming queries using append mode require that new rows are entirely independent of previously processed outputs, meaning they are simply appended as new files. Late data arriving after the watermark threshold is dropped entirely by Spark and never written to the Delta table, preventing unbounded state growth and maintaining strict correctness guarantees.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Late-arriving data arriving after the watermark is automatically updated in-place within the existing Delta files using ACID merge operations.
Why it's wrong here
Delta Lake append operations never perform in-place updates on historical files. New streaming micro-batches strictly write entirely new parquet files into the table directory structure without modifying existing data files, preserving transaction log integrity. Watermarked late data is dropped entirely rather than merged.
- ✗
Late-arriving data is buffered in an unlimited memory state store until the streaming query is manually stopped and restarted by an operator.
Why it's wrong here
Spark Structured Streaming never buffers state indefinitely without a bound. State stores rely on watermarks and TTL configurations to aggressively clean up old state data, preventing memory exhaustion and preventing unbounded growth during continuous operation.
- ✓
Late-arriving data that falls behind the specified watermark threshold is dropped and never written to the Delta table.
Why this is correct
Watermarks define how long the engine waits for late data. Any event whose event-time falls behind the current watermark is considered too late and is dropped to prevent unbounded state accumulation in streaming aggregations and joins.
- ✗
Late-arriving data forces the Delta Lake table to automatically switch its output mode to complete mode for that specific micro-batch.
Why it's wrong here
Output modes are statically defined in the streaming query definition via the DataFrameWriter API. An active streaming query cannot dynamically or automatically switch between append, update, and complete modes based on the arrival characteristics of incoming data.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.