DP-203 Develop data processing Practice Question
You are developing an Azure Databricks notebook that processes a large Delta Lake table. You must add a derived column that depends on the latest value of a watermark stored in a small reference table, and the notebook must refresh this value before each micro-batch. You need to ensure the reference data is re-read on every micro-batch rather than cached once. Which approach should you use?
⚠ Common exam trap
The trap here is assuming that caching or persisting the reference table will keep it current across micro-batches.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the foreachBatch sink and, inside the function, query the reference table with a fresh read on each invocation.
Structured Streaming captures static DataFrames when a query starts, so joins against them do not see later changes. foreachBatch executes a user-defined function for each micro-batch and permits arbitrary operations, including a fresh read of the reference table. Reading the watermark inside that function ensures the latest value is used for the derived column on every micro-batch, which is precisely the stated requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Persist the reference table with the MEMORY_AND_DISK storage level before starting the stream.
Why it's wrong here
Persisting a DataFrame caches its current contents and is the opposite of what is needed here. The cached snapshot would be reused across micro-batches, so any change to the watermark in the reference table would not be visible, and the derived column would be computed from outdated values.
- ✗
Broadcast the reference table as a static DataFrame and join it to the streaming DataFrame.
Why it's wrong here
Joining a static DataFrame to a streaming DataFrame is supported, but the static side is captured when the query starts. The watermark value would not be refreshed on subsequent micro-batches, so the derived column would use stale reference data and violate the requirement to re-read before each micro-batch.
- ✗
Enable the spark.databricks.delta.cache.enabled option on the streaming query.
Why it's wrong here
Delta caching accelerates repeated reads of the same files but does not introduce per-micro-batch refreshes. If the reference table changes, a cached read may still return old data, and the option does not alter how a streaming query captures static DataFrames, so the watermark would not be reliably updated.
- ✓
Use the foreachBatch sink and, inside the function, query the reference table with a fresh read on each invocation.
Why this is correct
foreachBatch gives you a Python or Scala function that runs once per micro-batch with the batch's DataFrame. Performing a fresh read of the reference table inside that function guarantees the watermark is retrieved anew for every micro-batch, which satisfies the refresh requirement while still allowing DataFrame operations on the batch.
Go deeper
Related to this question
About these practice questions
One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.