Databricks-DE-Assoc Data Transformation and Modeling Practice Question
A data engineering team is building a medallion architecture in Databricks. In the Silver layer, streaming data from Kafka must be cleaned, deduplicated, and written into a Delta table. Which Structured Streaming output mode should the engineer select to ensure append-only storage of fully processed, stateful deduplicated records?
⚠ Common exam trap
Candidates often mistakenly choose 'Complete' output mode when dealing with stateful deduplication, forgetting that complete mode requires rewriting the entire table on every trigger.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Append output mode
When performing stateful operations such as deduplication with dropDuplicates() on a streaming DataFrame, Structured Streaming requires the output mode to be set to append. This mode allows newly emitted unique rows to be written continuously to the target Delta table without requiring complete table rewrites, fitting the streaming append-only nature of the Silver medallion layer.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Complete output mode
Why it's wrong here
Complete output mode requires the entire updated result table to be written to the sink every time new data arrives. This mode is restricted to streaming aggregations where all rows are maintained in state and is incompatible with standard append-based Delta sinks used for raw or deduplicated streaming.
- ✗
Update output mode
Why it's wrong here
Update output mode writes only the rows that were updated in the result table since the last micro-batch. While useful for streaming queries with aggregations, it is not supported for streaming deduplication workflows writing to file-based Delta sinks without aggregations.
- ✓
Append output mode
Why this is correct
Append output mode is the default and mandatory output mode when executing stateful operations like dropDuplicates() on streaming DataFrames. It ensures that only new, non-duplicate rows resulting from each micro-batch are appended to the target Delta table sink.
- ✗
Overwrite output mode
Why it's wrong here
Overwrite rewrites the entire table on each micro-batch, destroying append-only history and breaking stateful deduplication across batches. It suits periodic full-refresh aggregations where the latest complete result replaces prior data, not streaming ingestion that must accumulate deduplicated records incrementally.
About these practice questions
One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.