Databricks-DE-Pro Developing Code (Python/SQL) Practice Question
A data engineer is implementing a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The engineer needs to ensure that the job can recover from failures without data loss or duplication. The job uses foreachBatch to perform upserts into the Delta table. Which checkpointing configuration is required to achieve exactly-once semantics?
⚠ Common exam trap
The trap here is assuming that any checkpoint location suffices, but a local path is not fault-tolerant, and append mode can duplicate data on replay.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the checkpointLocation to a reliable distributed storage path and ensure the Delta table write is idempotent using merge.
Exactly-once semantics in Structured Streaming require a reliable checkpoint location to track progress and idempotent writes to handle reprocessing. Using foreachBatch with MERGE provides idempotent upserts, and a distributed checkpoint location ensures recovery. The other options either use unreliable storage, non-idempotent writes, or unsupported write modes, failing to guarantee exactly-once.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Set the checkpointLocation to a reliable distributed storage path and ensure the Delta table write is idempotent using merge.
Why this is correct
Exactly-once semantics in Structured Streaming with foreachBatch require a reliable checkpoint location to track progress and idempotent writes to handle reprocessing. Using MERGE for upserts makes the write idempotent, so replays do not duplicate data. The checkpoint stores offsets and state, enabling recovery. This combination ensures exactly-once processing even after failures.
- ✗
Use the option 'startingOffsets' set to 'latest' and enable trigger once for the streaming query.
Why it's wrong here
Starting from latest offsets means the job will skip existing data, which is not about exactly-once but about initial position. Trigger once runs the query as a batch, but without proper checkpointing and idempotent writes, it does not guarantee exactly-once on failure. This option does not address recovery from failures.
- ✗
Set the checkpointLocation to a local path on the driver and use append mode for the Delta table write.
Why it's wrong here
A local checkpoint path is not reliable in a distributed environment; if the driver fails, the checkpoint is lost. Append mode does not provide idempotent writes, so reprocessing after failure would duplicate data. Thus, this configuration cannot guarantee exactly-once semantics and may lead to data loss or duplication.
- ✗
Set the Spark configuration spark.sql.streaming.checkpointLocation to a DBFS path and use complete mode for the Delta table write.
Why it's wrong here
While setting a checkpoint location is necessary, using complete mode for Delta table writes is not supported for streaming writes; Delta Lake only supports append and update modes. Complete mode would overwrite the entire table, which is not suitable for upserts and can lead to data loss. Thus, this option is incorrect.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.