Courseiva

Databricks-DE-Assoc Data Transformation and Modeling Practice Question

A data engineer has a Delta table named silver_events with columns event_id (string), event_ts (timestamp), and payload (string). The table is partitioned by event_date (derived from event_ts). The engineer needs to update the payload column for all events that occurred on '2024-06-01' based on a mapping table named updates (event_id, new_payload). Which PySpark operation should be used to perform this update efficiently while preserving Delta Lake ACID guarantees?

⚠ Common exam trap

The trap here is assuming that overwriting a partition or the entire table is equivalent to an update, but it lacks the atomicity and efficiency of MERGE INTO.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Perform a MERGE INTO operation on silver_events using updates as the source, matching on event_id and filtering both source and target to event_date = '2024-06-01'.

The correct approach is to use MERGE INTO, which is designed for upserts and updates based on a source table. Filtering both the target and source to the specific event_date partition ensures that only the relevant data is processed, optimizing performance. MERGE INTO also provides ACID guarantees, ensuring that the update is atomic and isolated, which is critical for maintaining data integrity in a Delta Lake table.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use DataFrame.write.mode('overwrite').partitionBy('event_date').saveAsTable('silver_events') with the updated DataFrame containing only the rows for '2024-06-01'.

    Why it's wrong here

    Overwriting with only the updated partition would replace the entire table or partition, potentially deleting all other partitions if not dynamically partitioned. Even with partitionBy, this approach is not atomic for partial updates and can lead to data loss or inconsistent state. It also does not leverage Delta Lake's ACID transactions for targeted updates, making it unsuitable for this scenario.

  • ✓

    Perform a MERGE INTO operation on silver_events using updates as the source, matching on event_id and filtering both source and target to event_date = '2024-06-01'.

    Why this is correct

    MERGE INTO is the correct Delta Lake operation for updating existing rows based on a source table. By matching on event_id and filtering both source and target to the specific event_date partition, the operation only scans and updates the relevant partition, which is efficient. It also preserves ACID guarantees, including atomicity and isolation, ensuring that concurrent readers see a consistent snapshot.

  • ✗

    Read the entire silver_events table into a DataFrame, apply a join with updates, and then use DataFrame.write.mode('overwrite').saveAsTable('silver_events') to persist the changes.

    Why it's wrong here

    This method rewrites the entire table, which is inefficient for updating a single partition and does not preserve ACID guarantees during the write. Concurrent readers might see an inconsistent state if the write fails midway. Additionally, it does not leverage Delta Lake's MERGE capabilities, leading to higher compute and storage costs, especially for large tables.

  • ✗

    Use the DeltaTable API's update method with a condition on event_date = '2024-06-01' and a join to the updates table.

    Why it's wrong here

    The DeltaTable API does not support joins in its update method; it only allows simple expressions. Attempting to join within an update would require a workaround that likely involves reading both tables and writing back, which is not efficient or ACID-compliant. The correct approach for join-based updates is MERGE INTO, not the update method.

About these practice questions

One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.