Courseiva
Develop data processing →hardMultiple Select

DP-203 Develop data processing Practice Question

You are implementing a Spark Structured Streaming job in Azure Databricks that consumes from an Azure Event Hubs topic and writes to a Delta table. The stream must tolerate reprocessing after a cluster restart without producing duplicate rows in the Delta table. You need to configure the write path accordingly. (Choose two.)

⚠ Common exam trap

The trap here is treating Delta Lake as if it enforced primary keys, when in fact it accepts duplicate rows unless the write logic itself is made idempotent.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable checkpointing by setting the checkpointLocation option on the writeStream.

Exactly-once behavior at the sink requires both a durable record of processed offsets and an idempotent write. Checkpointing supplies the first by persisting progress so a restart resumes correctly, and a MERGE keyed on a business identifier supplies the second by making replay harmless. Neither append mode, one-time triggers, nor file compaction can prevent logical duplicates because none of them compare row identity across batches.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable checkpointing by setting the checkpointLocation option on the writeStream.

    Why this is correct

    Checkpointing records the progress of each micro-batch, including which offsets have been processed, in durable storage. After a cluster restart, the stream resumes from the last committed offset rather than reprocessing the entire topic. Without a checkpoint location, Structured Streaming cannot track progress and will either fail to restart or re-read from the beginning, which is the primary cause of duplicate rows.

  • ✗

    Set the trigger to Trigger.Once so the stream reads the topic only one time.

    Why it's wrong here

    Trigger.Once processes all available data in a single batch and then stops the stream. It does not provide restart tolerance because the job terminates after one run, and any subsequent run would process new data without tracking prior offsets. It also does nothing to make the write idempotent, so re-running the job against the same offsets would duplicate rows.

  • ✗

    Enable auto compaction on the Delta table to merge small files and eliminate repeated rows.

    Why it's wrong here

    Auto compaction addresses the small-file problem by coalescing many small Parquet files into larger ones during writes. It optimizes file layout and read performance, but it never compares row contents or removes logical duplicates. Two identical rows written in separate files remain two rows after compaction, so this setting cannot satisfy the duplicate-avoidance requirement in the scenario.

  • ✗

    Configure the write mode as append and rely on the Delta table schema to reject duplicates.

    Why it's wrong here

    Append mode simply adds every row from each micro-batch to the Delta table. Delta does not enforce uniqueness on any column by default, so a replayed batch produces additional duplicate rows. Schema enforcement validates types and columns, not row identity, so this configuration cannot prevent the duplicates that the scenario requires you to avoid after a restart.

  • ✓

    Use foreachBatch with an idempotent MERGE INTO statement against the Delta table.

    Why this is correct

    foreachBatch lets you apply arbitrary logic to each micro-batch, and a MERGE INTO keyed on a stable business key makes the write idempotent. If a batch is replayed after a restart, the MERGE updates existing rows or inserts only new keys, so repeated processing does not create duplicates. Combined with checkpointing, this provides exactly-once semantics at the Delta table.

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.