DP-203 Develop data processing • Set 24
DP-203 Develop data processing Practice Test 24 — 15 questions with explanations. Free, no signup.
A data engineering team is building a batch processing solution for a financial services company. Data is ingested daily from multiple sources into Azure Data Lake Storage Gen2 in CSV format. The data must be transformed (filtered, aggregated, joined) and loaded into Azure Synapse Analytics dedicated SQL pool. The team must optimize for cost and performance. The total data volume is 2 TB per day. The team has the following options:
Option A: Use Azure Data Factory pipelines with copy activity to load raw CSV files into Synapse staging tables, then use T-SQL stored procedures in Synapse to perform transformations.
Option B: Use Azure Databricks with Auto Loader to incrementally ingest CSV files, perform transformations in Spark, and write the results to Synapse using the Spark Synapse connector.
Option C: Use Azure Data Factory with mapping data flows to transform the data in a serverless environment and then write to Synapse.
Option D: Use Azure Synapse Pipelines (built on ADF) with a notebook activity that runs a PySpark notebook in Synapse Spark pool to transform and load data.
Which option should the team choose to minimize cost and management overhead while meeting performance requirements?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.