Databricks-Spark-Assoc Structured Streaming Practice Question
A Databricks Structured Streaming job reads JSON files from a cloud storage directory and writes aggregated results to a Delta table. The source directory receives new files continuously and files are never modified after being written. The pipeline must tolerate late-arriving event data by up to 30 minutes and must not reprocess already-emitted windows when the query is restarted. Which combination of configurations is required to meet these requirements?
⚠ Common exam trap
The trap here is assuming that any watermark combined with any output mode will prevent reprocessing, when only append mode guarantees a window is emitted once after finalization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set a checkpoint location on a durable file system, apply `withWatermark("eventTime", "30 minutes")`, and use the `append` output mode with a Delta table sink.
The pipeline needs durable state recovery and late-data tolerance without re-emitting finalized windows. A checkpoint location on reliable storage ensures the query resumes from where it left off. A 30-minute watermark defines how long late events are accepted. The append output mode emits each window only once, after the watermark passes the window end, which matches the no-reprocessing requirement when writing to Delta.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure `spark.sql.streaming.checkpointLocation` to a reliable cloud storage path, use `withWatermark("eventTime", "30 minutes")`, and set the output mode to `update`.
Why it's wrong here
A persistent checkpoint location and a 30-minute watermark are correct for tolerating late data and maintaining state across restarts. However, the `update` output mode emits updated window results each time a late record arrives, which can cause duplicate emissions to the sink. This mode does not satisfy the requirement to avoid reprocessing already-emitted windows when writing to a non-idempotent sink.
- ✗
Set the source option `maxFilesPerTrigger` to a high value, define a watermark of 30 minutes on the event-time column, and rely on the default checkpoint location provided by the Spark session.
Why it's wrong here
Setting maxFilesPerTrigger only controls ingestion throughput and does not provide late-data tolerance. A watermark of 30 minutes is appropriate, but the default checkpoint location is not durable across restarts and can be deleted, causing the query to lose state and reprocess data. Without a persistent checkpoint, the no-reprocessing requirement cannot be guaranteed.
- ✓
Set a checkpoint location on a durable file system, apply `withWatermark("eventTime", "30 minutes")`, and use the `append` output mode with a Delta table sink.
Why this is correct
A durable checkpoint location preserves query progress and state across restarts, satisfying the no-reprocessing requirement. The 30-minute watermark allows the engine to accept and incorporate late events up to 30 minutes past the window boundary. The `append` output mode emits a window's final result only after the watermark passes the window end, ensuring each window is emitted exactly once to the Delta sink.
- ✗
Use `withWatermark("eventTime", "30 minutes")`, set the output mode to `complete`, and configure a checkpoint location on DBFS.
Why it's wrong here
The `complete` output mode rewrites the entire result table on every trigger, which is inefficient and can cause downstream consumers to see repeated full snapshots. While a checkpoint location and watermark are present, the complete mode does not provide the incremental, exactly-once emission of finalized windows that the scenario requires for a Delta sink without reprocessing.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.