Databricks-Spark-Assoc Structured Streaming Practice Question
You are building a Structured Streaming pipeline that reads from a Delta table source and applies a stateful deduplication using dropDuplicates on a composite key. After several hours, the job fails with an error indicating that the state store has grown too large. You need to bound the state size while still removing duplicate events that arrive within a reasonable window. Which approach should you take?
⚠ Common exam trap
The trap here is believing that tuning state store internals or partitioning can bound state, when only a watermark-based retention rule lets the engine decide a key can be safely forgotten.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add withWatermark on the event-time column and use dropDuplicatesWithinWatermark on the key columns.
Bounding state requires a rule that tells the engine when a key will no longer be needed. A watermark on event time provides that rule, and dropDuplicatesWithinWatermark uses it to evict keys older than the watermark. This keeps deduplication correct within the allowed lateness while preventing state from growing without limit. Configuration tweaks and repartitioning do not change the fundamental retention policy.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the checkpoint location to a new path and restart the query so the old state is discarded.
Why it's wrong here
Pointing to a fresh checkpoint location discards all prior state, which would reset deduplication memory but also lose progress tracking and any other stateful computation. The query would reprocess from the source's starting position depending on the source, potentially duplicating or skipping data, and the state would simply grow unbounded again. It is a destructive workaround, not a solution that bounds state while preserving correctness.
- ✗
Increase the spark.sql.streaming.stateStore.maintenanceInterval to force more frequent cleanup of old keys.
Why it's wrong here
The maintenance interval controls how often the state store performs internal cleanup tasks, but it does not define which keys are eligible for removal. With plain dropDuplicates and no watermark, all keys are retained indefinitely because the engine has no basis for deciding that a key will never be seen again. Tuning maintenance frequency may reduce fragmentation but cannot bound logical state size, so the query would still grow until it fails.
- ✓
Add withWatermark on the event-time column and use dropDuplicatesWithinWatermark on the key columns.
Why this is correct
dropDuplicatesWithinWatermark combined with withWatermark limits how long the engine retains keys for deduplication. Once the watermark passes a key's event time beyond the configured delay, the key can be evicted from state, bounding memory usage. This directly addresses unbounded state growth while still deduplicating events that arrive within the allowed lateness window, which matches the requirement to remove duplicates within a reasonable period.
- ✗
Repartition the stream by the composite key before calling dropDuplicates to spread state across more partitions.
Why it's wrong here
Repartitioning can distribute state across executors but does not reduce the total number of distinct keys retained. Each partition still stores every key it has seen forever, so aggregate state size continues to grow with the cardinality of the key space. This may delay the failure but does not bound state, and it adds shuffle overhead. The underlying need is a mechanism to expire keys, which repartitioning does not provide.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.