Databricks-Spark-Assoc Structured Streaming Practice Question
You are troubleshooting a Structured Streaming job that reads from Kafka and writes to a Delta table using foreachBatch. The job occasionally processes the same Kafka offsets twice after a task retry, causing duplicate rows in the Delta table. You want to ensure that each micro-batch's output is applied exactly once. Which change should you make inside the foreachBatch function?
⚠ Common exam trap
The trap here is assuming that Structured Streaming automatically guarantees exactly-once for arbitrary foreachBatch code, when in fact the custom function must be made idempotent by the developer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the batchId argument to make the Delta write idempotent, for example by merging on a batch identifier or using replaceWhere with the batchId.
Exactly-once sink semantics with foreachBatch require the batch function itself to be idempotent, because the function may be invoked more than once for the same batch. The provided batchId is stable across retries, so using it as a deduplication key or partition selector lets a retried write overwrite its previous output. Caching, append-only properties, and partition tuning do not provide this guarantee.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the batchId argument to make the Delta write idempotent, for example by merging on a batch identifier or using replaceWhere with the batchId.
Why this is correct
foreachBatch provides the batch identifier, which is monotonically increasing and stable across retries of the same batch. By using it to make the write idempotent, such as merging on a batch column or replacing a partition keyed by batchId, a retried batch overwrites rather than duplicates its prior output. This delivers exactly-once effects at the sink even when the batch function is re-executed.
- ✗
Increase the number of shuffle partitions so each batch writes to more files.
Why it's wrong here
Shuffle partition count affects parallelism and file layout, not correctness of repeated writes. More partitions may improve throughput or reduce file sizes, but a retried batch would still append the same rows again. This option does not introduce any deduplication key or transactional boundary, so it cannot prevent duplicates after a task retry.
- ✗
Call df.cache() on the batch DataFrame before writing to Delta.
Why it's wrong here
Caching materializes the batch DataFrame in memory to avoid recomputation within a single execution, but it does not persist across retries or provide any deduplication at the sink. If the batch function is invoked again after a failure, the cached data is gone and the write repeats. Caching is a performance optimization, not an idempotency mechanism, so duplicates would still occur.
- ✗
Set the Delta table property delta.appendOnly to true.
Why it's wrong here
delta.appendOnly disallows updates and deletes on the table, enforcing append-only semantics. That is the opposite of what is needed here, because idempotent application typically requires merges or partition replacement, which are updates. Enabling this property would prevent the corrective write and cause failures, and it does not address duplicate application of a retried batch.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.