Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A structured streaming DataFrame writes to a Delta table with a foreachBatch function that performs an upsert. After a cluster restart, the stream reprocesses some micro-batches and duplicate rows appear in the target table. The foreachBatch code already uses MERGE keyed on a unique id. Which change best prevents duplicates after restart?
⚠ Common exam trap
The trap here is believing a MERGE on a unique key alone guarantees exactly-once, when restart reprocessing is actually governed by the streaming checkpoint.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a checkpointLocation to the writeStream options so the stream records its progress and resumes from the last committed offset.
Structured streaming tracks processed offsets and state in the checkpoint location. Without a checkpoint, a restarted query cannot know what it already committed and replays earlier micro-batches, creating duplicates even with an idempotent MERGE. Supplying checkpointLocation lets the query resume from the last committed offset, so each batch is processed once. Combined with the existing MERGE on a unique id, this yields the intended exactly-once behavior on restart.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the trigger to processingTime='1 minute' so micro-batches are larger and less likely to overlap.
Why it's wrong here
Trigger interval controls how often micro-batches are generated, not whether committed progress is remembered. A longer interval changes batch size and latency but does not create a durable record of processed offsets. After restart the query would still reprocess from the start without a checkpoint. Trigger tuning is unrelated to the duplicate-row root cause.
- ✓
Add a checkpointLocation to the writeStream options so the stream records its progress and resumes from the last committed offset.
Why this is correct
Structured streaming relies on the checkpoint location to persist progress, offsets, and state. Without it, a restarted query has no record of which micro-batches were committed and reprocesses data, producing duplicates even when the MERGE key is unique. Setting checkpointLocation gives exactly-once semantics for the sink when combined with an idempotent operation such as MERGE. This is the direct fix for reprocessing after restart.
- ✗
Switch the output mode from append to complete so each micro-batch is fully replaced in the target table.
Why it's wrong here
Complete mode rewrites the entire result table for every micro-batch, which is meant for aggregations and would not preserve the upsert semantics of the MERGE. It also increases write cost dramatically and still depends on checkpointing to know its progress. Changing output mode does not address the missing checkpoint. It would not prevent duplicates after a restart.
- ✗
Add a deduplication step using dropDuplicates on the unique id inside the foreachBatch before the MERGE.
Why it's wrong here
dropDuplicates removes duplicates within a single micro-batch, but the duplicates here come from reprocessing across restarts, so each batch looks unique on its own. It also discards legitimate repeated keys that should update. The MERGE already handles key uniqueness in the target. This treats a symptom without restoring stream progress tracking.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.