Courseiva
Develop data processing →hardMultiple Choice

DP-203 Develop data processing Practice Question

You are implementing a data processing solution in Azure Databricks. The solution reads JSON files from Azure Data Lake Storage Gen2, performs complex transformations using PySpark, and writes the results to a Delta table. You need to ensure that the write operation is idempotent and can recover from failures without duplicating data. Which approach should you use?

⚠ Common exam trap

The trap here is assuming that append mode with post-deduplication is sufficient for idempotence, when merge is the designed mechanism for upserts.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the merge operation with a unique key to upsert records into the Delta table.

The merge operation in Delta Lake enables upserts based on a unique key, ensuring that rerunning the job after a failure does not duplicate data. It provides ACID transactions and idempotence, which are critical for reliable data processing. Append mode and overwrite mode do not offer the same guarantees, and temporary files with Copy activity lack transactional integrity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the merge operation with a unique key to upsert records into the Delta table.

    Why this is correct

    Delta Lake's merge operation allows you to upsert records based on a unique key, making the write idempotent. If the job fails and is rerun, the merge will update existing records and insert new ones without creating duplicates. This satisfies the requirement for idempotent and recoverable writes.

  • ✗

    Write the output in append mode and use a deduplication step after each run.

    Why it's wrong here

    Append mode adds new rows, and while a subsequent deduplication can remove duplicates, it is not inherently idempotent and requires additional processing. If the job fails mid-write, the deduplication may not run, leaving duplicates. This approach is less reliable than a merge operation for ensuring idempotence.

  • ✗

    Write the output using the overwrite mode to replace the entire Delta table on each run.

    Why it's wrong here

    Overwrite mode replaces the entire table, which is not idempotent for incremental loads and can cause data loss if the job fails before completion. It also does not support recovery from failures without duplicating or losing data. This approach is unsuitable for the requirement of idempotent writes with failure recovery.

  • ✗

    Write the output to a temporary Parquet file and then use a Copy activity to move it to the Delta table.

    Why it's wrong here

    This approach does not provide idempotence because the Copy activity would append or overwrite the destination, and failure recovery would require manual intervention. It also bypasses Delta Lake's transactional guarantees. Using merge directly is simpler and more reliable for idempotent writes.

About these practice questions

Courseiva writes every DP-203 question from scratch — 509 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.