DP-203 Develop data processing Practice Question
You are designing a data processing solution for a retail company that uses Azure Synapse Analytics. The solution must process point-of-sale (POS) data from multiple stores. The data arrives in CSV files in Azure Data Lake Storage Gen2. Each store sends a file every hour. You need to process the files as they arrive and load the data into a dedicated SQL pool. The solution must handle late-arriving files (files that arrive after the scheduled processing time) and ensure that the data is consistent. Which approach should you use?
⚠ Common exam trap
DP-203 often tests the choice between different data loading patterns, where candidates might overlook the need for upserts and late-arriving data handling, opting for simpler append-only methods like PolyBase CTAS, which do not ensure consistency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Azure Data Factory with a Copy activity to load data into a staging table in the dedicated SQL pool, then use a Stored Procedure activity to merge the data into the final table.
The recommended approach for handling late-arriving files and ensuring data consistency when loading into a dedicated SQL pool is to use Azure Data Factory with a Copy activity to load data into a staging table, then use a Stored Procedure activity to merge the data into the final table. This pattern allows for upserts and handles late-arriving data by merging based on business keys.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Azure Data Factory with a Copy activity to load data into a staging table, then use a Data Flow activity to perform upserts.
Why it's wrong here
A scheduled Copy activity plus Data Flow upserts processes on a trigger cadence, so late-arriving files wait for the next run and consistency lags. It is tempting because Data Factory orchestrates batch pipelines well, but the stem demands event-driven processing as files arrive.
- ✗
Use Azure Databricks to read the CSV files, perform upserts, and write to the dedicated SQL pool using JDBC.
Why it's wrong here
Databricks with JDBC writes bypasses Synapse's native ingestion and adds cluster overhead, and it does not inherently trigger on file arrival. It is tempting because Databricks handles complex transformations and upserts, but the stem requires event-driven loading into a dedicated SQL pool.
- ✗
Use PolyBase to create external tables on the CSV files and then use CREATE TABLE AS SELECT to load into the dedicated SQL pool.
Why it's wrong here
PolyBase external tables read the CSV files at query time and CREATE TABLE AS SELECT performs a one-off bulk copy; neither detects newly arrived hourly files nor reprocesses late-arriving ones, so consistency across stores breaks. It suits scheduled bulk ingestion of a static file set, not continuous arrival handling.
- ✓
Use Azure Data Factory with a Copy activity to load data into a staging table in the dedicated SQL pool, then use a Stored Procedure activity to merge the data into the final table.
Why this is correct
Staging then merging via a Stored Procedure activity makes the load idempotent: late-arriving files are upserted on business keys rather than duplicated. This satisfies both the hourly arrival pattern and the consistency requirement in the dedicated SQL pool.
Go deeper
Related to this question
Learn chapter
Implement Azure Synapse Analytics
Key term
Azure Synapse Analytics
Azure Synapse Analytics is a cloud-based data integration, warehousing, and analytics service that brings together big data and data warehouse capabilities under one platform.
Key term
Azure Data Factory
Azure Data Factory is a cloud-based data integration service that lets you create, schedule, and orchestrate data pipelines to move and transform data from various sources to destinations.
About these practice questions
Courseiva writes every DP-203 question from scratch — 509 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.