Courseiva
Develop data processing →mediumMultiple Choice

DP-203 Develop data processing Practice Question

A data engineering team is building a batch processing solution for a financial services company. Data is ingested daily from multiple sources into Azure Data Lake Storage Gen2 in CSV format. The data must be transformed (filtered, aggregated, joined) and loaded into Azure Synapse Analytics dedicated SQL pool. The team must optimize for cost and performance. The total data volume is 2 TB per day. The team has the following options:

Option A: Use Azure Data Factory pipelines with copy activity to load raw CSV files into Synapse staging tables, then use T-SQL stored procedures in Synapse to perform transformations.

Option B: Use Azure Databricks with Auto Loader to incrementally ingest CSV files, perform transformations in Spark, and write the results to Synapse using the Spark Synapse connector.

Option C: Use Azure Data Factory with mapping data flows to transform the data in a serverless environment and then write to Synapse.

Option D: Use Azure Synapse Pipelines (built on ADF) with a notebook activity that runs a PySpark notebook in Synapse Spark pool to transform and load data.

Which option should the team choose to minimize cost and management overhead while meeting performance requirements?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Option C

Correct answer: B (Option C — Azure Data Factory mapping data flows). Mapping data flows execute on a serverless Azure Data Factory integration runtime, scaling automatically and costing only per run, which minimizes cost and management overhead. Answer choice A (Option B, Azure Databricks with Auto Loader) requires managing Spark clusters. Answer choice C (Option A, Azure Data Factory copy activity plus T-SQL stored procedures) requires staging tables and stored procedures. Answer choice D (Option D, Synapse Pipelines with a notebook activity) requires managing Synapse Spark pools. Therefore, mapping data flows are the most cost-effective and lowest-overhead option.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Option B

    Why it's wrong here

    Requires cluster management and higher cost.

  • ✓

    Option C

    Why this is correct

    Serverless, cost-effective, low overhead.

  • ✗

    Option A

    Why it's wrong here

    Requires manual staging and T-SQL management.

  • ✗

    Option D

    Why it's wrong here

    Requires Synapse Spark pool management.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

Go deeper

Related to this question

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.