Courseiva

DP-203 · domain

Develop data processing

Develop data processing (46%) covers batch and streaming pipelines in Azure. Expect questions on Azure Databricks, Synapse Spark and serverless SQL pools, Data Factory, Stream Analytics, and Event Hubs. You must choose correct transformations, authentication, windowing, and file formats for ADLS Gen2 and Delta Lake scenarios.

185 questions59 easy75 medium51 hard

Focused practice

Practice Develop data processing questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Develop data processing

You must design and implement batch and streaming transformations using Databricks, Synapse, Data Factory, and Stream Analytics. The most important thing is selecting the right compute engine, authentication method, and windowing or file format for each ADLS Gen2 scenario.

Joining large Parquet datasets in Azure Databricks with Delta Lake and partitioning

Querying Parquet in Synapse serverless SQL pools using OPENROWSET and managed identity

Transforming semi-structured JSON with Synapse Spark or serverless SQL JSON functions

Streaming aggregation in Stream Analytics with Event Hubs, tumbling windows, and late-arrival tolerance

Watch out for

Common Develop data processing exam traps

  • ▸Assuming serverless SQL pools can use SQL authentication or service principal secrets instead of Azure AD pass-through or managed identity.
  • ▸Forgetting that Stream Analytics late-arrival tolerance affects event ordering and window output, not just dashboard refresh.
  • ▸Using row-by-row processing in Databricks instead of set-based DataFrame operations, causing poor performance on large Parquet joins.

Question index

All Develop data processing questions (185)

Click any question to see the full explanation, or start a practice session above.

1

Your company runs a streaming job in Azure Stream Analytics that ingests data from Event Hubs and outputs to Azure Synapse Analytics. The job is failing with a 'Watermark delay' alert and the output to Synapse is delayed by over 30 minutes. The input rate is 5,000 events per second. The job uses a 1-minute tumbling window. What is the most likely cause of the delay?

Hard
2

You have an Azure Data Factory pipeline that must copy data from an on-premises Oracle database to Azure Blob Storage every night. The on-premises server cannot accept inbound connections, and no VPN or ExpressRoute is available. You need to enable connectivity with minimal administrative overhead. What should you deploy?

Easy
3

You are implementing a data processing solution in Azure Databricks. The solution must read data from Azure Data Lake Storage Gen2, transform it using PySpark, and write the results back to a different location in the same storage account. You need to authenticate to the storage account securely without storing secrets in the notebook. What should you use?

Easy
4

You have a mission-critical pipeline that processes financial transactions in Azure Synapse Analytics. The pipeline uses Azure Data Factory with a mapping data flow to transform data. You need to ensure high availability and minimal data loss in case of a regional failure. What should you implement?

Hard
5

You have a streaming pipeline using Azure Stream Analytics that ingests data from Event Hubs and outputs to Azure Synapse Analytics. The job has a high watermark delay and is falling behind. You need to reduce the latency. Which action should you take?

Hard
6

You are optimizing the performance of a large-scale batch processing job in Azure Databricks. The job reads data from Azure Data Lake Storage Gen2, performs transformations, and writes results back. You notice that the job is I/O bound. Which THREE strategies can improve performance? (Choose three.)

Hard
7

You are tasked with transforming data in an Azure Synapse Analytics pipeline using a mapping data flow. The source data contains a column 'FullName' in the format 'LastName, FirstName'. You need to split this into two separate columns: 'LastName' and 'FirstName'. Which transformation should you use?

Easy
8

You are designing a data processing solution in Azure Synapse Analytics. The solution must process streaming data from Azure Event Hubs and store the results in a dedicated SQL pool. You need to choose the most appropriate service for near real-time ingestion with minimal latency. What should you use?

Medium
9

You are designing an ETL process in Azure Data Factory. You need to transform data using Mapping Data Flows. Which THREE of the following transformations are available in Mapping Data Flows?

Medium
10

You are troubleshooting a slow-running pipeline in Azure Data Factory that uses a Copy activity to transfer data from Azure Blob Storage to Azure Synapse Analytics. The pipeline processes about 100 GB of CSV files. The copy performance is poor even though the source and sink are in the same region. What is the most likely cause?

Medium
11

You are running a Python script in Azure Databricks that reads a CSV file from DBFS. The script runs successfully in an interactive notebook but fails when executed as a job with the error: 'Path does not exist: dbfs:/tmp/data.csv'. What is the most likely cause?

Easy
12

You have an Azure Databricks notebook that processes a large Delta table. The notebook uses a structured streaming query to read from the Delta table and write to another Delta table. The source table receives frequent updates and deletes. You need the streaming query to process both new data and changes (updates and deletes) from the source table. What should you do?

Hard
13

You are designing a data processing solution for a retail company that uses Azure Synapse Analytics. The solution must process point-of-sale (POS) data from multiple stores. The data arrives in CSV files in Azure Data Lake Storage Gen2. Each store sends a file every hour. You need to process the files as they arrive and load the data into a dedicated SQL pool. The solution must handle late-arriving files (files that arrive after the scheduled processing time) and ensure that the data is consistent. Which approach should you use?

Hard
14

Your organization uses Azure Synapse Analytics serverless SQL pool to query Parquet files in Azure Data Lake Storage Gen2. You notice that queries are slow when filtering on a date column. You need to improve query performance without increasing costs. What should you do?

Hard
15

You are developing an Azure Databricks notebook that processes streaming data from Azure Event Hubs using Structured Streaming. The stream writes to a Delta Lake table. You need to ensure that the stream can recover from failures and continue processing from where it left off without reprocessing all data. You also need to minimize the impact on the source. What should you configure?

Medium
16

Refer to the exhibit. You are deploying an Azure Synapse Analytics dedicated SQL pool using the provided ARM template snippet. After deployment, you need to adjust the performance level to DW200c to handle increased workload. Which parameter should you modify?

Hard
17

You are monitoring an Azure Stream Analytics job that processes data from an IoT hub. The job's output to Azure Synapse Analytics is experiencing high latency. The job's SU% utilization is at 90%. Which action will most likely reduce the latency?

Hard
18

You are building a data pipeline that uses Azure Data Factory to copy data from a REST API to Azure Blob Storage. The REST API returns JSON data in pages of 1000 records each. The total number of records is 50,000. Which activity or feature should you use to loop through the pages?

Medium
19

Refer to the exhibit. You are deploying an Azure Synapse Analytics workspace using an ARM template. The template defines a managed virtual network integration runtime. You need to ensure that the integration runtime can run mapping data flows with a time-to-live (TTL) of 10 minutes. What is the purpose of the 'timeToLive' property in this configuration?

Medium
20

You are developing an Azure Synapse Analytics serverless SQL pool solution that queries Parquet files in Azure Data Lake Storage Gen2. Analysts run ad-hoc queries with predicates on a high-cardinality column named TransactionId, and each query scans the entire folder, causing high cost. You need to reduce the amount of data scanned per query without changing the file format. What should you do?

Hard
21

Which TWO actions can you take to optimize the performance of a dedicated SQL pool in Azure Synapse Analytics when loading large volumes of data?

Medium
22

Refer to the exhibit. You have a mapping data flow in Azure Data Factory that aggregates sales data. The data flow runs successfully but the sink table contains only the total sum per run instead of per product. What is missing?

Easy
23

You are building a batch processing solution in Azure Synapse Analytics that reads data from a dedicated SQL pool, applies complex transformations using Synapse Spark, and writes the results back to the dedicated SQL pool. The pipeline must run on a schedule and handle transient failures with retries. Which approach should you use?

Hard
24

You are partitioning a large fact table in Azure Synapse Dedicated SQL Pool by date. The table is used for queries that filter on CustomerID and Date. You want to minimize data movement. Which distribution strategy should you use?

Medium
25

You have an Azure Synapse Analytics workspace with a dedicated SQL pool. You need to create an external table that references Parquet files stored in Azure Data Lake Storage Gen2. The external table will be used for ad-hoc queries. Which statement correctly describes the required components?

Medium
26

You need to transform JSON data containing nested arrays into a tabular format for analysis in Azure Synapse Analytics. Which transformation in Azure Data Factory or Synapse Pipelines should you use?

Easy
27

Refer to the exhibit. You are reviewing an Azure Stream Analytics job query. The job has a stream input and a reference data input. The job is failing with the error 'Reference data input must be of type Reference, not Stream'. What is the cause of the error?

Easy
28

You are designing a data processing solution in Azure using Azure Data Lake Storage Gen2 as the storage layer. You need to ensure that data ingested from various sources is immutable and can be used for both batch and streaming workloads. Which storage design pattern should you implement?

Medium
29

You are developing an Azure Databricks notebook that processes a large Delta Lake table. You must add a derived column that depends on the latest value of a watermark stored in a small reference table, and the notebook must refresh this value before each micro-batch. You need to ensure the reference data is re-read on every micro-batch rather than cached once. Which approach should you use?

Medium
30

Your organization uses Azure Data Lake Storage Gen2 (ADLS Gen2) and wants to transform data using Azure Databricks. The data is stored in Parquet format. You need to read the data into a Spark DataFrame. Which DataFrame reader method should you use?

Easy
31

You are creating an Azure Data Factory pipeline that must copy data from an on-premises SQL Server to Azure Blob Storage daily. The on-premises network restricts inbound connections, and you need a secure connection without exposing the SQL Server to the public internet. What should you use to connect to the on-premises SQL Server?

Easy
32

You need to process a large dataset that contains personally identifiable information (PII). The data must be anonymized before being used for analytics. Which Azure service should you use to apply column-level masking dynamically?

Easy
33

You are using Azure Synapse Analytics to process data in a dedicated SQL pool. You need to ensure that queries against a large fact table perform well. The fact table is partitioned by date and distributed by a product key. Which two actions should you take? (Choose two.)

Medium
34

You are developing a data processing pipeline for a gaming company that uses Azure Databricks. The pipeline processes game event data from Azure Event Hubs. You need to detect cheating patterns by analyzing events in real time. The solution must be able to handle high throughput and low latency. The output should be written to Azure Cosmos DB for real-time dashboards. Which approach should you use?

Medium
35

You are developing an Azure Databricks notebook to process streaming data from Azure Event Hubs. The notebook must write the processed data to a Delta table with exactly-once processing guarantees. You need to configure the write operation. Which option should you use?

Easy
36

You are designing a data processing solution using Azure Databricks. You need to read data from Azure Data Lake Storage Gen2, transform it using Spark SQL, and write to a Delta table. Which TWO configurations are required to ensure optimal performance for large datasets?

Medium
37

You need to process streaming data from Azure Event Hubs and store the results in Azure Cosmos DB for a real-time dashboard. The solution must handle duplicate events and ensure exactly-once processing. Which Azure service should you use?

Easy
38

You are designing a data processing solution for a retail company that uses Azure Databricks. The solution needs to process streaming sales data from Event Hubs and batch data from Azure Data Lake Storage Gen2. You need to ensure that the solution can handle late-arriving data and maintain exactly-once semantics. Which TWO technologies should you use?

Hard
39

You have an Azure Synapse Analytics dedicated SQL pool. A nightly ELT process loads a 500 GB staging table and then applies transformations using a stored procedure. The procedure performs many single-row updates against a large fact table, and the load now exceeds its window. You need to reduce the duration of the transformation step. What should you do?

Hard
40

You maintain an Azure Stream Analytics job that reads from an Event Hubs input and writes to an Azure Synapse Analytics dedicated SQL pool. During peak load the job produces late-arriving events that are dropped, and downstream reports show missing rows. You must retain and process events that arrive after the watermark by up to several minutes without changing the input. What should you configure?

Hard
41

You are designing a batch processing solution for a data lake. Source files arrive daily in Parquet format in Azure Data Lake Storage Gen2. The data must be cleaned, aggregated, and loaded into an Azure Synapse SQL pool. The solution should minimize compute costs and management overhead. Which technology should you use for the transformation?

Easy
42

You are a data engineer for a healthcare company that processes patient data. You have an Azure Databricks workspace with a cluster configured for data processing. You need to implement a solution that processes streaming data from Azure Event Hubs, enriches it with reference data stored in Azure Cosmos DB, and writes the output to Delta Lake in Azure Data Lake Storage Gen2. The solution must ensure that the data processing is fault-tolerant and can handle schema evolution. The reference data is updated infrequently. You need to choose an approach that minimizes complexity and cost. What should you do?

Hard
43

You are designing a streaming data solution for IoT devices that generate 10,000 events per second. The data must be processed with sub-second latency and then stored in Azure Data Lake Storage Gen2 for archival. Which Azure service should you use for the stream processing?

Medium
44

You are designing a data processing solution for an e-commerce company that uses Azure Synapse Analytics. The solution must process clickstream data from a web application. The data arrives in JSON format through Azure Event Hubs. You need to load the data into a dedicated SQL pool every 5 minutes with minimal latency. The data volume is about 100 MB every 5 minutes. You want to use PolyBase for loading. Which approach should you use?

Medium
45

You are optimizing a pipeline in Azure Data Factory that copies data from Azure Blob Storage to Azure Synapse Analytics. The pipeline uses a copy activity with PolyBase. The data is partitioned by date in Blob Storage. You notice that the load is slow. What is the most likely cause?

Hard
46

You are developing an Azure Databricks notebook that reads a large Delta table, performs a join with a smaller reference table, and writes the result back to Delta Lake. The job runs on a cluster with autoscaling enabled and frequently spills to disk during the join. You need to reduce shuffle and improve performance without changing the result. Which action should you take?

Hard
47

Which TWO Azure services can be used to perform real-time data processing on streaming data?

Medium
48

Which TWO techniques can you use to handle schema drift in Azure Data Factory mapping data flows?

Easy
49

Refer to the exhibit. You have an Azure Data Factory pipeline that copies data from a CSV file in Blob Storage to a Synapse dedicated SQL pool table named dbo.Sales. The pipeline fails. The error message indicates that the 'Amount' column in the sink table does not allow NULLs but the source contains NULL values. What is the best way to resolve this issue without losing data?

Hard
50

You are implementing a mapping data flow in Azure Data Factory that joins a large fact table in Azure Synapse Analytics with a slowly changing dimension (SCD) table in Azure SQL Database. The fact table has 500 million rows and the dimension has 2 million rows. You need to optimize the join performance and minimize data movement. The dimension table is small enough to fit in memory. Which join type should you configure in the data flow?

Hard
51

You are building an Azure Data Factory pipeline that calls an external REST API returning a JSON array of records. The API paginates results using a 'nextLink' field in the response body, and the number of pages varies per run. You must ingest all pages into Azure Blob Storage in a single pipeline run. Which activity configuration should you use?

Medium
52

You are building an Azure Stream Analytics job that reads from an Azure Event Hub capturing device telemetry. The job must emit results into an Azure Synapse Analytics dedicated SQL pool. You need to minimize latency and avoid intermediate storage. What should you do?

Medium
53

You are building an Azure Synapse Analytics pipeline that processes JSON files landing in Azure Data Lake Storage Gen2. The files contain nested arrays representing order line items. You need to flatten this nested structure into a tabular format within a Mapping Data Flow before loading to a dedicated SQL pool. The solution must minimize data movement and avoid writing intermediate files to storage. Which transformation should you use to flatten the nested arrays?

Medium
54

Which THREE options are valid ways to transform data in Azure Synapse Analytics?

Medium
55

You are designing a data processing solution in Azure Synapse Analytics. The solution must process streaming data from Azure Event Hubs and store the results in a dedicated SQL pool. The solution must support exactly-once semantics and handle late-arriving data. Which Azure service should you use to implement this solution?

Medium
56

You are building an Azure Stream Analytics job that processes JSON telemetry from Azure Event Hubs. The events contain a nested array field named `readings` with sensor values. You need to transform the data so that each sensor reading becomes a separate output row, and then write the results to Azure Synapse Analytics. Which two actions should you perform? (Choose two.)

Medium
57

You are implementing a medallion architecture in Azure Databricks. The silver layer must contain deduplicated, conformed records, and the gold layer must serve aggregated reporting tables. You need to choose Delta Lake operations that support incremental, idempotent updates as new bronze files arrive. Which two operations should you use? (Choose two.)

Hard
58

You are using Azure Data Lake Storage Gen2 as the data lake for your organization. You need to process files in the 'incoming' folder using a scheduled Azure Databricks notebook. After processing, the files should be moved to the 'processed' folder. The files are large (up to 10 GB) and you want to minimize the time to move them. Which approach should you use?

Medium
59

You are implementing a real-time analytics solution using Azure Stream Analytics. The job ingests data from Azure Event Hubs and must output to an Azure SQL Database. You need to ensure that the job can handle out-of-order events and produce accurate aggregations over 5-minute windows. Which setting should you configure?

Hard
60

Your team runs Azure Data Factory pipelines that must copy files from an on-premises file share to Azure Data Lake Storage Gen2 on a nightly schedule. The on-premises network blocks inbound connections and the data must not be exposed to the public internet. You need to enable connectivity without opening firewall ports. What should you deploy?

Easy
61

You are creating an Azure Synapse Analytics pipeline that must copy data from an Azure SQL Database into a dedicated SQL pool. The destination table already exists and the pipeline must append new rows without truncating existing data. Which staging and load option should you configure in the Copy activity?

Easy
62

You are analyzing a Kusto query in Azure Data Explorer that calculates total sales per product for January 2024 and filters for products with sales over 10,000. The query uses the materialize() function. You notice that the query runs slower than expected. What is the primary reason the materialize() function may not be providing the expected performance benefit in this query?

Hard
63

You are optimizing a Spark DataFrame transformation in Azure Synapse Analytics. The DataFrame has 20 columns and 100 million rows. You notice that the job is slow due to many small files being written to the output. Which two actions can you take to reduce the number of output files? (Choose two.)

Easy
64

You are building an Azure Stream Analytics job that reads JSON telemetry from an Azure Event Hub, calculates a 5-minute tumbling window average per device, and writes results to an Azure Synapse Analytics dedicated SQL pool. The stream must handle occasional bursts of late-arriving events by including events that arrive up to 3 minutes after the window closes. You need to configure the job's event ordering settings to meet the late-arrival requirement while minimizing memory usage. What should you do?

Medium
65

You are designing a data transformation solution for a retail company. The company receives daily CSV files from 200 stores via SFTP. The files must be cleaned, validated, and aggregated before loading into Azure Synapse dedicated SQL pool. The solution must minimize administrative overhead and support easy monitoring. Which approach do you recommend?

Medium
66

Which TWO of the following are supported sources for Azure Data Factory Copy activity? (Choose two.)

Easy
67

You have a dedicated SQL pool in Azure Synapse that stores a fact table with over 100 billion rows. Query performance is degrading over time. You notice that the table is hash-distributed on a column with many duplicate values. What is the most likely impact?

Medium
68

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline uses a self-hosted integration runtime and runs successfully during business hours. However, after a recent network security update, the pipeline fails with a connection error to the on-premises SQL Server. What is the most likely cause?

Easy
69

You are optimizing a data pipeline in Azure Synapse Analytics that loads data from a CSV file in ADLS Gen2 into a dedicated SQL pool using PolyBase. The load is slow and you need to improve performance. Which action would be MOST effective?

Hard
70

You are implementing a mapping data flow in Azure Data Factory that processes data from an Azure SQL Database. The data flow includes a derived column transformation that adds a new column based on a complex expression. You need to ensure the expression handles null values appropriately. Which function should you use to replace null values with a default?

Medium
71

Your team uses Azure Synapse Analytics serverless SQL pool to query Parquet files in Azure Data Lake Storage Gen2. The query performance is inconsistent, and some queries take a long time to execute. You need to improve query performance. What should you do?

Hard
72

A data engineering team is building a batch processing solution for a financial services company. Data is ingested daily from multiple sources into Azure Data Lake Storage Gen2 in CSV format. The data must be transformed (filtered, aggregated, joined) and loaded into Azure Synapse Analytics dedicated SQL pool. The team must optimize for cost and performance. The total data volume is 2 TB per day. The team has the following options: Option A: Use Azure Data Factory pipelines with copy activity to load raw CSV files into Synapse staging tables, then use T-SQL stored procedures in Synapse to perform transformations. Option B: Use Azure Databricks with Auto Loader to incrementally ingest CSV files, perform transformations in Spark, and write the results to Synapse using the Spark Synapse connector. Option C: Use Azure Data Factory with mapping data flows to transform the data in a serverless environment and then write to Synapse. Option D: Use Azure Synapse Pipelines (built on ADF) with a notebook activity that runs a PySpark notebook in Synapse Spark pool to transform and load data. Which option should the team choose to minimize cost and management overhead while meeting performance requirements?

Medium
73

You are designing a real-time analytics solution for IoT devices that emit telemetry data every second. The data must be aggregated every minute and stored in Azure SQL Database for historical analysis. You need to minimize latency and operational overhead. Which approach should you recommend?

Hard
74

Your organization is using Azure Synapse Analytics dedicated SQL pool. You notice that queries are running slower than expected. Upon reviewing the execution plans, you see that some queries are performing table scans instead of seeks on large fact tables. What is the most likely cause?

Medium
75

You are developing an Azure Data Factory pipeline that must call an external REST API, parse the JSON response, and load selected fields into an Azure SQL Database. The API requires a bearer token that expires every hour, so the pipeline must obtain a fresh token before each call. You need to implement the token acquisition and header injection without writing custom code in a data flow. What should you use?

Easy
76

You are implementing a Spark Structured Streaming job in Azure Databricks that reads from an Azure Event Hubs topic. The job must handle late-arriving data up to 10 minutes and produce aggregated results every 5 minutes. You need to configure the watermark and window. Which code snippet should you use?

Hard
77

Your team is developing a data processing solution in Azure Synapse Analytics. You need to ensure that the solution can automatically scale compute resources based on workload demand for serverless SQL pools. Which feature should you configure?

Easy
78

You need to incrementally load new and updated records from a source SQL Server database to Azure Synapse Dedicated SQL Pool. The source table has a LastModifiedDate column. Which Azure Data Factory feature should you use to implement incremental loading efficiently?

Easy
79

You are designing a data pipeline in Azure Synapse Analytics to ingest data from Azure Blob Storage into a dedicated SQL pool. The source files are CSV with varying row lengths, and you need to ensure optimal performance for reads. Which file format and compression should you recommend?

Medium
80

You are creating an Azure Data Factory data flow that transforms data from an Azure SQL Database. The data flow must filter rows based on a column value and then aggregate the results. Which transformation should you use first?

Easy
81

You are developing a data processing solution that requires aggregating sales data from multiple CSV files stored in Azure Data Lake Storage Gen2. The data should be cleansed and transformed before loading into Azure Synapse Analytics. Which Azure service should you use to implement a code-free transformation pipeline?

Easy
82

You are implementing a solution in Azure Databricks that reads from a Delta Lake table and writes to another Delta Lake table. The pipeline must process data incrementally and handle updates and deletes from the source. Which feature should you use to read the changes?

Hard
83

You are implementing a Mapping Data Flow in Azure Synapse Analytics that reads from a Parquet source and writes to a Delta sink. The data flow includes a Surrogate Key transformation to generate unique keys for each row. You notice that when the data flow runs multiple times, the surrogate keys are not consistent across runs and sometimes overlap. You need to ensure that surrogate keys are unique and stable across runs. What should you do?

Hard
84

Which TWO of the following are required components to set up a data pipeline that uses Change Data Capture (CDC) to incrementally load data from SQL Server to Azure Synapse using Azure Data Factory?

Easy
85

You are building a real-time dashboard to monitor user activity on a website. The data is ingested via Azure Event Hubs and must be aggregated every minute with a 30-second late-arrival tolerance. The aggregated results should be stored in Azure Cosmos DB for low-latency reads. Which Azure service should you use to perform the windowed aggregation?

Medium
86

You are designing a data processing solution in Azure Data Factory that uses mapping data flows. You need to perform type conversions on incoming data. Which two transformations can be used to change data types? (Choose two.)

Easy
87

You are designing a data processing solution for a financial services company. The solution must process sensitive customer data and comply with GDPR. The data will be stored in Azure Synapse Analytics. You need to ensure that only authorized users can view specific columns (e.g., credit card numbers). Which security feature should you implement?

Hard
88

You are building an Azure Stream Analytics job that reads from an Azure Event Hubs input and writes to an Azure Synapse Analytics dedicated SQL pool. You need to compute a 5-minute tumbling window aggregation that outputs only once per window after all events for that window have arrived. Which query construct should you use?

Medium
89

You are developing an Azure Databricks notebook that processes streaming data from Azure Event Hubs and writes to a Delta Lake table. The stream must handle late-arriving data up to 30 minutes old and ensure that aggregations are computed correctly even if events arrive out of order. You need to minimize state store size and avoid unbounded growth. Which combination of features should you use?

Hard
90

You are designing a data processing solution in Azure Synapse Analytics. The solution must use a dedicated SQL pool to store fact and dimension tables. The fact table is expected to have billions of rows. Which distribution strategy should you recommend for the fact table to optimize query performance and minimize data movement?

Easy
91

Refer to the exhibit. You are creating a serverless SQL table in Azure Synapse Analytics that reads Parquet files from the specified location. The folder contains multiple Parquet files with different schemas. When querying the table, you get an error about schema mismatch. What is the most likely reason?

Hard
92

You are implementing a Spark Structured Streaming job in Azure Databricks that consumes from an Azure Event Hubs topic and writes to a Delta table. The stream must tolerate reprocessing after a cluster restart without producing duplicate rows in the Delta table. You need to configure the write path accordingly. (Choose two.)

Hard
93

You are designing a data pipeline that uses Azure Data Factory to load data from an FTP server to Azure Data Lake Storage. The FTP server requires authentication with username and password. Which type of linked service should you create?

Easy
94

You are creating an Azure Data Factory pipeline that must copy data from an on-premises Oracle database to Azure Blob Storage every night. The on-premises network restricts inbound connections. You need to configure the integration runtime. What should you do?

Easy
95

You are authoring an Azure Databricks notebook that reads Parquet files from Azure Data Lake Storage Gen2 and must write results to a Delta table. Users report that queries against the Delta table return stale data after each notebook run, even though the write succeeds. You need to ensure readers always see the latest committed data. What should you do?

Medium
96

You are running a pipeline in Azure Data Factory that uses a Mapping Data Flow. The data flow reads from Azure SQL Database and writes to Azure Synapse Analytics. You find that the data flow is very slow. Which configuration change would most likely improve performance?

Medium
97

You are using Azure Databricks to process a large dataset stored in Delta Lake. You need to reduce the number of files scanned during queries by organizing data into folders based on a commonly filtered column. Which Delta Lake feature should you implement?

Easy
98

You are designing a data processing solution that uses Azure Databricks to transform large datasets. You need to ensure that the processing is cost-effective and can scale to handle variable workloads. Which cluster configuration should you recommend?

Medium
99

You are designing a data processing solution for a marketing company that uses Azure Synapse Analytics. The solution needs to process customer data from multiple sources, including CRM and web analytics. The data must be cleansed and transformed before loading into a dedicated SQL pool. The transformations include string manipulations, date conversions, and lookups. You need to choose a serverless transformation approach that integrates with Azure Synapse pipelines. Which approach should you use?

Easy
100

You are building an Azure Stream Analytics job that reads JSON events from an Azure Event Hub and writes aggregated results to an Azure Synapse Analytics dedicated SQL pool. The events include a field named `EventTime` that is sometimes missing or malformed. You need the job to process only events with a valid `EventTime` and route invalid events to a separate output for later inspection. What should you do?

Medium
101

A company uses Azure Synapse Analytics dedicated SQL pool. The data engineering team notices that queries against a large fact table are running slowly. The table uses round-robin distribution and has a columnstore index. The team wants to improve query performance without adding more resources. Which action should the team take?

Medium
102

You have a pipeline in Azure Data Factory that copies data from on-premises SQL Server to Azure Blob Storage. The pipeline fails with a 'Connection timed out' error. You have already verified that the Integration Runtime is running and the SQL Server firewall allows connections from the Integration Runtime. What should you check next?

Easy
103

You are designing a near-real-time data processing solution for a retail company. The source is a Kafka cluster on-premises. The target is an Azure Synapse Dedicated SQL Pool. The solution must handle up to 10,000 events per second with less than 5-minute latency. Which Azure service should you use to ingest the data?

Hard
104

You are designing a data processing solution in Azure Synapse Analytics. The solution must support both batch and streaming data ingestion. Which Azure service should you use to ingest streaming data into Synapse Analytics?

Easy
105

You are building an Azure Data Factory pipeline that processes files from Azure Blob Storage. The pipeline uses a Mapping Data Flow to transform the data and then writes the output to Azure Data Lake Storage Gen2. You need to ensure that the Data Flow can handle schema drift, where incoming files may have additional columns not present in the initial schema. What should you configure in the Data Flow?

Medium
106

You are designing a data processing solution for a global company. Data must be processed in near real-time and aggregated by region. You need to minimize latency for downstream consumers. Which Azure service should you use for stream processing?

Medium
107

You are developing an Azure Stream Analytics job that ingests telemetry from Azure Event Hubs and writes results to an Azure Synapse Analytics dedicated SQL pool. The job must compute a 5-minute tumbling window aggregation and write the aggregated rows to the dedicated SQL pool. You need to configure the output so that each window's aggregated rows are written efficiently. What should you do?

Medium
108

Your company uses Azure Synapse Analytics to run a large-scale batch processing job every night. The job currently runs on a dedicated SQL pool and takes 4 hours. Management wants to reduce the runtime to under 2 hours without increasing cost. The job involves heavy compute operations with no data movement limitations. What should you do?

Hard
109

Refer to the exhibit. You have an Azure Data Factory pipeline that performs an incremental load from an Azure SQL Database source to a target Azure SQL Database. The pipeline uses a watermark column approach. After running the pipeline, you notice that the target table is empty. What is the most likely cause of this issue?

Hard
110

You have an Azure Databricks notebook that processes a large Delta table and must be orchestrated from Azure Data Factory on a schedule. The notebook accepts two parameters, the source path and a run date. You need the pipeline to pass these values at runtime and to surface notebook failures as pipeline failures. Which activity configuration should you use?

Medium
111

Which THREE of the following are best practices for optimizing performance of Delta Lake tables in Azure Synapse Analytics? (Choose three.)

Hard
112

You have an Azure Databricks notebook that processes data from a Delta table. The notebook runs slowly due to many small files. You need to optimize the Delta table for faster reads. Which Delta Lake operation should you run?

Easy
113

You are building a data processing pipeline in Azure Synapse Analytics. The pipeline should read data from Azure Data Lake Storage Gen2 (Parquet files), apply transformations using a mapping data flow, and write the results to a dedicated SQL pool table. The source data contains personally identifiable information (PII). You need to mask the PII columns (e.g., email) using a data masking function within the data flow. Which transformation should you use?

Hard
114

You are building an Azure Synapse Analytics pipeline that uses a Mapping Data Flow to transform data from Azure Data Lake Storage Gen2. The data flow includes a derived column transformation that uses a custom expression to calculate a new field. You need to debug the data flow and preview the output at each transformation. Which feature should you use?

Medium
115

You are building a real-time dashboard that displays sales data from an Azure SQL Database. The dashboard must refresh every 30 seconds with minimal latency. You need to choose the appropriate Azure service for data processing and visualization. Which service should you use?

Medium
116

You are designing an Azure Synapse Analytics pipeline that uses a Mapping Data Flow to transform data from Azure Data Lake Storage Gen2. The data flow must handle schema drift, where new columns can appear in the source files over time. You need to ensure that the new columns are automatically included in the sink output without modifying the data flow. What should you do?

Medium
117

You are building an Azure Data Factory pipeline that must copy data from an on-premises SQL Server to Azure Blob Storage. The pipeline runs on a schedule every hour. You need to ensure that the copy activity can securely access the on-premises SQL Server. What should you configure?

Easy
118

You are using Azure Data Factory to copy data from an Azure SQL Database to an Azure Data Lake Storage Gen2 account. The copy activity is failing intermittently with a timeout error. You need to improve the throughput and reliability of the copy operation. What should you do?

Medium
119

You are processing a large dataset in Azure Synapse Analytics using a dedicated SQL pool. You need to load data from Parquet files in Azure Data Lake Storage Gen2 into a staging table, then transform and load into a fact table. The fact table is partitioned by date. You want to maximize query performance and minimize data movement. Which technique should you use?

Hard
120

You are designing a streaming solution in Azure Synapse Analytics using the serverless SQL pool to query streaming data in real-time. The data is ingested via Azure Event Hubs and processed using Azure Stream Analytics. The output of Stream Analytics is written to Azure Data Lake Storage Gen2 in Delta Lake format. You need to ensure that the serverless SQL pool can query the latest data with minimal latency. Which approach should you use?

Medium
121

You need to transform semi-structured JSON data into a tabular format for analysis in Azure Synapse Analytics. The data is stored in ADLS Gen2. Which feature should you use to query the JSON data directly without loading it into a table?

Easy
122

You are developing an Azure Data Factory pipeline that processes data from an on-premises SQL Server. The pipeline uses a self-hosted integration runtime. You need to ensure that the pipeline can handle schema changes in the source table without failing. What should you do?

Hard
123

You are designing a batch processing solution in Azure Synapse Analytics using pipelines. The solution must load data from multiple sources (Azure Blob Storage, Azure SQL Database, and REST API) into a dedicated SQL pool. After loading, you need to run a stored procedure to aggregate the data. Which two activities should you include in the pipeline? (Choose two.)

Medium
124

You are developing an Azure Databricks notebook that processes streaming data from Azure Event Hubs. The notebook must write the processed data to a Delta table. You need to ensure that the stream can handle late data and update previously written records. Which Delta Lake feature should you use?

Medium
125

You are designing a batch processing solution using Azure Databricks. The data source is a large Parquet dataset stored in Azure Data Lake Storage Gen2 (ADLS Gen2). The processing requires joining two datasets: one with 10 billion rows and another with 1 million rows. The cluster uses Photon runtime. Which optimization should you apply to minimize shuffle?

Hard
126

You are building an Azure Stream Analytics job that reads JSON events from an Azure Event Hub. Each event contains a nested array property named 'readings' with multiple sensor values. You need to output one row per sensor reading to an Azure Synapse Analytics dedicated SQL pool. The query must flatten the array. Which query syntax should you use?

Medium
127

Which TWO are benefits of using Azure Databricks Auto Loader for incremental data ingestion?

Easy
128

Which TWO Azure services can be used to perform data transformation in a data pipeline? (Select two.)

Easy
129

Which THREE components are required to implement a real-time data processing solution using Azure Stream Analytics?

Hard
130

You are designing a data processing solution for an e-commerce company. The company receives millions of clickstream events per hour from their website and needs to aggregate the data by product category and windowed time intervals for real-time dashboards. You need to minimize latency and cost. Which service should you use?

Medium
131

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline uses a self-hosted integration runtime. You notice that the copy activity fails intermittently with the error: 'Failure happened on 'Source' side. ErrorCode=SqlOperationFailed'. The on-premises SQL Server is under heavy load during business hours. What is the most likely cause?

Hard
132

You are designing a data processing pipeline in Azure Data Factory. The pipeline must copy data from Azure Blob Storage to Azure SQL Database and transform the data using a mapping data flow. The data flow includes a Derived Column transformation. What is the purpose of the Derived Column transformation?

Easy
133

You are implementing a data processing solution in Azure Databricks. The solution reads JSON files from Azure Data Lake Storage Gen2, performs complex transformations using PySpark, and writes the results to a Delta table. You need to ensure that the write operation is idempotent and can recover from failures without duplicating data. Which approach should you use?

Hard
134

You are implementing a Spark Structured Streaming job in Azure Databricks that reads from an Azure Event Hubs topic and writes to a Delta table. The job must handle late-arriving data up to 10 minutes and aggregate counts per device every 5 minutes. Which combination of settings should you use?

Hard
135

You need to orchestrate a data pipeline that includes a Python script and a Data Flow in Azure Synapse Analytics. The Python script must run before the Data Flow. Which activity should you use to run the Python script?

Easy
136

You are using Azure Synapse Analytics dedicated SQL pool to process large fact tables. You need to improve query performance for joins between a large fact table and a small dimension table. The dimension table is less than 2 GB. What should you do?

Medium
137

You are running a Spark job in Azure Synapse Analytics that reads from a Delta Lake table and performs multiple transformations. The job fails with an out-of-memory error on the executors. Which action should you take first to resolve the issue?

Easy
138

Your team is developing a data processing solution that uses Azure Databricks to transform streaming data from Azure Event Hubs. The transformation includes joining the stream with a static reference table stored in Azure Data Lake Storage Gen2. You need to implement the join efficiently. Which approach should you use?

Easy
139

Refer to the exhibit. You have an Azure Synapse Analytics workspace. You need to ensure that data processing jobs can access the Data Lake Storage Gen2 account using a managed identity. What should you do?

Hard
140

You are using Azure Databricks to process a large dataset stored in Azure Data Lake Storage Gen2. The data is in Parquet format and you need to optimize read performance for a query that filters on a specific column. What should you do?

Easy
141

You are developing an Azure Databricks notebook that processes JSON files stored in Azure Data Lake Storage Gen2. You need to read the files into a DataFrame and automatically infer the schema. Which code should you use?

Easy
142

You are designing a streaming job in Azure Stream Analytics. The job needs to count the number of events per device type every 10 seconds. The input is from Event Hubs. Which query should you use?

Easy
143

You are using Azure Synapse Analytics dedicated SQL pool to run a query that joins a large fact table (10 billion rows) and a small dimension table (1 million rows). The query is slow. Which distribution strategy should you use for the dimension table to improve performance?

Medium
144

A manufacturing company uses Azure Data Lake Storage Gen2 to store IoT sensor data. The data arrives in JSON format with a nested structure. You need to transform the data into a tabular format for downstream analytics using Azure Synapse Pipelines. Which data flow transformation should you use?

Medium
145

You are implementing a data processing solution in Azure Synapse Analytics using Spark pools. The solution reads Parquet files from Azure Data Lake Storage Gen2, performs transformations, and writes the results to a dedicated SQL pool. You need to optimize the write performance to the dedicated SQL pool. Which technique should you use?

Medium
146

You are monitoring an Azure Data Factory pipeline that copies data from Azure Blob Storage to Azure SQL Database. The pipeline fails intermittently with the error: 'Operation on target SQL table failed: String or binary data would be truncated.' Which action should you take to resolve this issue?

Easy
147

You are a data engineer at a manufacturing company. You need to process sensor data from IoT devices that arrive in real time. The data is sent to Azure Event Hubs. You need to aggregate the data over 5-minute windows and store the results in Azure Data Lake Storage Gen2 in Parquet format. The solution should minimize cost and use serverless components. Which solution should you use?

Medium
148

You are a data engineer at a financial services company. You are developing a data processing pipeline that uses Azure Data Factory to copy transactional data from an Azure SQL Database to Azure Data Lake Storage Gen2. The pipeline runs daily and processes about 10 GB of data. You need to implement error handling for the pipeline. Specifically, if the copy activity fails due to a transient error, the pipeline should retry automatically. If the retry fails, the pipeline should log the error and send an email alert to the operations team. What should you do?

Easy
149

You are designing a data processing pipeline that ingests data from a REST API endpoint every hour. The API returns JSON data with a varying schema. You need to store the raw data in Azure Data Lake Storage Gen2 and later process it using Azure Databricks. Which file format should you use for the raw data storage?

Medium
150

You are writing a T-SQL query against a dedicated SQL pool in Azure Synapse Analytics. The query aggregates a fact table containing billions of rows by joining it to a small dimension table. You observe that the join produces a large amount of data movement and the query runs slowly. You need to reduce data movement for this recurring pattern. What should you do?

Medium
151

You are reviewing a Spark job definition in Azure Synapse Analytics. The job aggregates sales data. The job runs successfully but takes longer than expected. You notice that dynamic allocation is disabled and the executor instances are fixed at 10. The cluster has a maximum of 20 nodes. What is the most likely reason for the slow performance?

Medium
152

Which TWO options are correct approaches to handle schema drift in Azure Data Factory Mapping Data Flows?

Medium
153

You have an Azure Data Factory pipeline that executes a stored procedure in Azure SQL Database. The pipeline fails with an error indicating that the stored procedure ran out of memory. What change should you make to the pipeline to resolve this?

Medium
154

You are developing a data processing solution in Azure Synapse Analytics. The solution must use a serverless SQL pool to query Parquet files stored in Azure Data Lake Storage Gen2. Which authentication method should you use to ensure that the queries use the identity of the caller and adhere to Azure role-based access control (RBAC) permissions?

Easy
155

Your organization uses Azure Synapse Analytics dedicated SQL pool to store sales data. You need to design a data loading process for a nightly batch that inserts new rows and updates existing rows based on the business key. The table has a clustered columnstore index. Which approach minimizes table fragmentation?

Medium
156

You are processing streaming data from IoT devices using Azure Stream Analytics. The data includes temperature readings and device IDs. You need to calculate the average temperature per device over a 5-minute window, sliding every 1 minute. Which window function should you use?

Easy
157

You are developing a real-time data processing solution using Azure Stream Analytics. The input is an Azure Event Hubs stream with JSON data containing a 'timestamp' field. You need to output the average temperature per device every minute using a tumbling window. Which query should you use?

Easy
158

Refer to the exhibit. You submit a Spark job in Azure Synapse Analytics using the Azure CLI. The job runs slowly during the shuffle phase. The input data is about 200 GB. Which configuration change would best improve performance for this shuffle-heavy workload?

Hard
159

You are using Azure Synapse Analytics serverless SQL pool to query Parquet files in Azure Data Lake Storage Gen2. The query returns fewer rows than expected. What should you check first?

Easy
160

You need to transform data in Azure Databricks using Apache Spark. The data is stored in Delta Lake format in Azure Data Lake Storage Gen2. Which method should you use to read the data into a Spark DataFrame?

Easy
161

Which TWO actions can you take to optimize query performance in Azure Synapse Analytics dedicated SQL pool?

Medium
162

You are designing a data processing solution using Azure Databricks with Delta Lake. You need to ensure ACID transactions and schema enforcement. Which feature should you enable?

Hard
163

Which TWO components are required to set up a streaming data pipeline using Azure Synapse Analytics? (Select two.)

Medium
164

You are using Azure Synapse Analytics to process streaming data from Azure Event Hubs. The data must be written to a Delta Lake table in ADLS Gen2 with exactly-once semantics. Which processing engine should you use?

Easy
165

Which TWO features of Azure Databricks help manage data governance and security for sensitive data?

Easy
166

You are building a streaming pipeline in Azure Stream Analytics that reads JSON events from an Azure Event Hub and writes to Azure Synapse Analytics. The events include a nested array of sensor readings. You need to flatten this array so each reading becomes a separate row. Which Stream Analytics feature should you use?

Medium
167

You have a production pipeline in Azure Data Factory that copies data from an on-premises SQL Server to Azure Blob Storage using a self-hosted integration runtime. The pipeline fails intermittently with a 'Connection closed' error. The data volume is 50 GB per run. What should you first troubleshoot to resolve this issue?

Hard
168

You are optimizing an Azure Synapse Analytics dedicated SQL pool that processes large fact tables. You need to improve query performance for a common join between a fact table and a dimension table. The fact table is distributed using hash distribution on a column that is not the join key. The dimension table is small and replicated. You want to minimize data movement during the join. What should you do?

Hard
169

You are troubleshooting a failed Azure Synapse Pipeline execution. The pipeline uses a Copy activity to load data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The error indicates a 'Connection timeout' to the on-premises source. The Integration Runtime is Self-Hosted and has been running successfully for months. What is the most likely cause?

Medium
170

You have a Data Factory pipeline that runs a U-SQL script in Azure Data Lake Analytics. The script processes terabytes of data and outputs to a CSV file. The pipeline is failing with the error: 'The job failed with UserError: Script execution failed.' You need to troubleshoot the issue. Which approach should you take first?

Hard
171

You are designing a data processing solution for a healthcare organization. The solution must process streaming data from IoT devices and store it in Azure Data Lake Storage Gen2. The data must be available for both real-time dashboards and historical analysis. You need to minimize operational overhead. What should you do?

Hard
172

Your organization uses Azure Synapse Analytics. You need to design a data transformation pipeline that processes streaming data from Azure Event Hubs, performs aggregations over a 5-minute tumbling window, and loads the results into a dedicated SQL pool table. Which Azure service should you use to implement the streaming transformation?

Medium
173

You are building an Azure Stream Analytics job that ingests telemetry from Azure Event Hubs and writes aggregated results to an Azure Synapse Analytics dedicated SQL pool. The job must compute a five-minute tumbling window average per device and tolerate events that arrive up to three minutes late. During testing, you observe that events arriving after the window closes are silently dropped. You need to ensure late events are included in the correct window result. What should you configure in the Stream Analytics job?

Medium
174

You are designing a data lakehouse architecture in Azure using Delta Lake. The solution needs to process batch and streaming data from multiple sources, including IoT devices and CRM systems. You need to ensure data quality by enforcing schema validation and handling schema evolution. You also need to provide a unified catalog for querying. Which service should you use?

Medium
175

You are designing a data processing solution in Azure Databricks to transform streaming data from Azure Event Hubs. The data must be aggregated in 1-minute tumbling windows and written to Azure Synapse Analytics. Which Spark API should you use?

Easy
176

Which TWO are valid ways to process data in Azure Synapse Analytics?

Medium
177

You need to perform incremental data loading from Azure SQL Database to Azure Data Lake Storage Gen2 using Azure Data Factory. Which approach is the most efficient?

Easy
178

You are designing a data processing solution in Azure Data Factory that must process files as they arrive in Azure Blob Storage. The solution must trigger a pipeline automatically when a new file is created, and then run a Databricks notebook to process the file. You need to configure the trigger and the pipeline activity. Which two actions should you perform? (Choose two.)

Medium
179

You have an Azure Synapse Analytics dedicated SQL pool. A nightly ELT process loads data into a staging table using PolyBase, then transforms and inserts it into a large fact table. You need to minimize data movement during the transformation step and ensure the fact table is optimized for large range scans. Which table distribution and index should you choose for the fact table?

Medium
180

You are implementing a streaming pipeline in Azure Stream Analytics that reads from an Azure Event Hub and writes aggregated results to an Azure Synapse Analytics dedicated SQL pool. The query groups events into 30-second windows. You need to ensure that the job can handle late-arriving events up to 2 minutes after the window closes without dropping them. What should you configure?

Medium
181

You are designing a data pipeline to ingest streaming data from IoT devices into Azure Synapse Analytics. The data must be available for querying with minimal latency, but you also need to handle spikes in throughput without data loss. Which service should you use as the ingestion layer?

Medium
182

Refer to the exhibit. You have an Azure Data Factory pipeline that copies trade data from Azure Blob Storage to Azure SQL Database. The pipeline runs every hour and truncates the destination table before each copy. However, users report that data is missing during the copy window. What is the most likely cause?

Medium
183

Your team is developing a data processing solution using Azure Databricks. The data is stored in Delta Lake format in Azure Data Lake Storage Gen2. You need to ensure that when multiple jobs concurrently write to the same Delta table, the operations are atomic and consistent. Which Delta Lake feature should you use?

Easy
184

You are designing a data processing solution in Azure Databricks. The data is stored in Azure Data Lake Storage Gen2 and you need to perform transformations using Apache Spark. The security requirements mandate that all data in transit must be encrypted and that the storage account must not be accessible from the public internet. What should you configure?

Easy
185

You need to transform data in Azure Synapse Analytics using a language that supports procedural logic and error handling. Which option should you use?

Easy

Frequently asked questions

What does the Develop data processing domain cover on the DP-203 exam?
You must design and implement batch and streaming transformations using Databricks, Synapse, Data Factory, and Stream Analytics. The most important thing is selecting the right compute engine, authentication method, and windowing or file format for each ADLS Gen2 scenario.
How many questions are in this domain?
This page lists all 185 Develop data processing questions in the DP-203 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Develop data processing questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
dp-203 DP-203 develop data processing Practice Questions