Courseiva
Structured Streaming →mediumMultiple Choice

Databricks-Spark-Assoc Structured Streaming Practice Question

An engineer is developing a Structured Streaming job that reads from an Apache Kafka source and writes the output continuously toDelta Lake using outputMode("append"). The stream occasionally experiences late-arriving data. Which downstream behavior can the engineer expect regarding the Delta Lake table?

⚠ Common exam trap

Candidates often assume that late-arriving data is automatically updated in the destination table or buffered indefinitely, forgetting that watermarks strictly drop data that falls behind the threshold in append mode.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Late-arriving data that falls behind the specified watermark threshold is dropped and never written to the Delta table.

Streaming queries using append mode require that new rows are entirely independent of previously processed outputs, meaning they are simply appended as new files. Late data arriving after the watermark threshold is dropped entirely by Spark and never written to the Delta table, preventing unbounded state growth and maintaining strict correctness guarantees.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Late-arriving data arriving after the watermark is automatically updated in-place within the existing Delta files using ACID merge operations.

    Why it's wrong here

    Delta Lake append operations never perform in-place updates on historical files. New streaming micro-batches strictly write entirely new parquet files into the table directory structure without modifying existing data files, preserving transaction log integrity. Watermarked late data is dropped entirely rather than merged.

  • ✗

    Late-arriving data is buffered in an unlimited memory state store until the streaming query is manually stopped and restarted by an operator.

    Why it's wrong here

    Spark Structured Streaming never buffers state indefinitely without a bound. State stores rely on watermarks and TTL configurations to aggressively clean up old state data, preventing memory exhaustion and preventing unbounded growth during continuous operation.

  • ✓

    Late-arriving data that falls behind the specified watermark threshold is dropped and never written to the Delta table.

    Why this is correct

    Watermarks define how long the engine waits for late data. Any event whose event-time falls behind the current watermark is considered too late and is dropped to prevent unbounded state accumulation in streaming aggregations and joins.

  • ✗

    Late-arriving data forces the Delta Lake table to automatically switch its output mode to complete mode for that specific micro-batch.

    Why it's wrong here

    Output modes are statically defined in the streaming query definition via the DataFrameWriter API. An active streaming query cannot dynamically or automatically switch between append, update, and complete modes based on the arrival characteristics of incoming data.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.