Databricks-DE-Assoc Data Transformation and Modeling Practice Question
Exhibit
{
"operation": "MERGE INTO target USING source ON target.id = source.id",
"when_matched": "UPDATE SET target.val = source.val",
"when_not_matched": "INSERT *",
"error": "java.lang.IllegalArgumentException: Multiple source rows matched one target row"
}Refer to the exhibit. The merge operation is failing in your pipeline. What is the root cause of this error, and how should it be resolved?
⚠ Common exam trap
Candidates often attempt to resolve MERGE errors by changing the join condition or adding more columns to the ON clause, rather than addressing the root cause: duplicate source keys.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The source DataFrame must be aggregated or deduplicated to ensure one match per target ID.
The error occurs because the join condition in the MERGE statement is not unique relative to the source data. When the source contains multiple records with the same identifier that exists in the target, the engine cannot decide which source record to use for the update. To resolve this, you must deduplicate the source data before performing the merge, ensuring each key has exactly one entry.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The target table does not have a primary key defined.
Why it's wrong here
Delta Lake does not enforce primary key constraints for MERGE operations in the same way relational databases do. The issue is logically rooted in the source data's cardinality, not the absence of a formal key constraint in the table metadata itself. This error is purely about source ambiguity.
- ✓
The source DataFrame must be aggregated or deduplicated to ensure one match per target ID.
Why this is correct
This is the classic resolution for 'multiple matches'. By using a window function or 'dropDuplicates' on the source DataFrame, you ensure that for every ID, there is only one source record. This removes the ambiguity that causes the MERGE operation to fail during the execution phase.
- ✗
The target table is too large to perform a MERGE and should be replaced with a full overwrite.
Why it's wrong here
While full overwrites are an option, they are inefficient and often logistically impossible for large tables. The error is a logic error, not a capacity error. Replacing a MERGE with a full overwrite would be an expensive and unnecessary workaround that does not solve the underlying data quality issue.
- ✗
The cluster configuration needs more memory to handle the join operation.
Why it's wrong here
The error is an IllegalArgumentException, which indicates a violation of the MERGE operation's logic, not a resource constraint. Increasing memory will not change the fact that the source data is structurally incorrect for the requested MERGE operation. The source data itself must be prepared properly before the join.
Visual reference
About these practice questions
One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.