DP-203 Develop data processing Practice Question
You are building a batch processing solution in Azure Synapse Analytics that reads data from a dedicated SQL pool, applies complex transformations using Synapse Spark, and writes the results back to the dedicated SQL pool. The pipeline must run on a schedule and handle transient failures with retries. Which approach should you use?
⚠ Common exam trap
Many candidates confuse Azure Batch or Azure Functions as viable alternatives for Spark job orchestration, overlooking the native integration and retry capabilities of Synapse Pipelines within the same service.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Azure Synapse Pipelines with a Notebook activity that runs Spark code
Azure Synapse Pipelines with a Notebook activity is the correct approach because it natively integrates Synapse Spark for complex transformations and supports scheduling and retry policies for transient failures. This allows you to read from a dedicated SQL pool, process data in Spark, and write back to the pool without external orchestration, leveraging the built-in pipeline reliability features.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Azure Batch with a custom application to run Spark jobs
Why it's wrong here
Azure Batch schedules arbitrary executables on pools of VMs; it provides no native Spark execution engine, no Synapse integration, and no built-in retry semantics for Spark stages. It is tempting as a generic compute scheduler, and would be correct for embarrassingly parallel custom binaries, not for orchestrating Synapse Spark transformations.
- ✓
Use Azure Synapse Pipelines with a Notebook activity that runs Spark code
Why this is correct
A Notebook activity inside Azure Synapse Pipelines runs the Spark transformation code while the surrounding pipeline supplies scheduling and built-in retry policies for transient failures. This satisfies both the batch orchestration and resilience requirements without external tooling.
- ✗
Use Azure Functions to trigger Spark jobs on demand
Why it's wrong here
Azure Functions triggers jobs on demand, providing no schedule, dependency orchestration or built-in retry policy for transient failures. It suits event-driven, short-lived compute. A scheduled pipeline with retry handling belongs in Synapse pipelines or Azure Data Factory with Spark activities.
- ✗
Use Azure Databricks with Auto Loader and Delta Live Tables
Why it's wrong here
Auto Loader and Delta Live Tables ingest files into Delta Lake, not a dedicated SQL pool, and Databricks sits outside the Synapse workspace the scenario specifies. It is tempting because Delta Live Tables does schedule pipelines with automatic retries, but for file-based ingestion into Delta, not Synapse Spark reading and writing a dedicated SQL pool.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.