Courseiva

DP-203 · domain

troubleshooting

Practise Microsoft Azure Data Engineer Associate DP-203 troubleshooting practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

509 questions165 easy211 medium133 hard

Focused practice

Practice troubleshooting questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about troubleshooting

troubleshooting questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common troubleshooting exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All troubleshooting questions (509)

Click any question to see the full explanation, or start a practice session above.

1

You need to store structured data in Azure. The data will be accessed by multiple applications using T-SQL queries, and you require automatic indexing and serverless compute. Which Azure service should you use?

Easy
2

You are securing an Azure Data Lake Storage Gen2 account that contains sensitive data. Which TWO of the following should you implement to protect data from unauthorized access?

Medium
3

Your company runs a streaming job in Azure Stream Analytics that ingests data from Event Hubs and outputs to Azure Synapse Analytics. The job is failing with a 'Watermark delay' alert and the output to Synapse is delayed by over 30 minutes. The input rate is 5,000 events per second. The job uses a 1-minute tumbling window. What is the most likely cause of the delay?

Hard
4

You have an Azure Data Factory pipeline that must copy data from an on-premises Oracle database to Azure Blob Storage every night. The on-premises server cannot accept inbound connections, and no VPN or ExpressRoute is available. You need to enable connectivity with minimal administrative overhead. What should you deploy?

Easy
5

You are implementing a data processing solution in Azure Databricks. The solution must read data from Azure Data Lake Storage Gen2, transform it using PySpark, and write the results back to a different location in the same storage account. You need to authenticate to the storage account securely without storing secrets in the notebook. What should you use?

Easy
6

Your company uses Azure Data Lake Storage Gen2 and needs to implement a data retention policy that automatically deletes files older than 90 days in a specific container. What should you use?

Medium
7

You are a data engineer at a financial services company. You have an Azure Data Lake Storage Gen2 account named finlake that stores sensitive transaction data in Parquet files. You need to ensure that data is encrypted at rest using a customer-managed key stored in Azure Key Vault, and that the key is automatically rotated every 90 days. You also need to be able to revoke access to the data immediately if the key is compromised. What should you do?

Medium
8

You need to secure data at rest for an Azure Data Lake Storage Gen2 account that contains sensitive financial data. Which configuration should you enable to ensure that data is encrypted using a customer-managed key stored in Azure Key Vault, and that access to the key is logged?

Easy
9

You have a mission-critical pipeline that processes financial transactions in Azure Synapse Analytics. The pipeline uses Azure Data Factory with a mapping data flow to transform data. You need to ensure high availability and minimal data loss in case of a regional failure. What should you implement?

Hard
10

You have a streaming pipeline using Azure Stream Analytics that ingests data from Event Hubs and outputs to Azure Synapse Analytics. The job has a high watermark delay and is falling behind. You need to reduce the latency. Which action should you take?

Hard
11

You are optimizing the performance of a large-scale batch processing job in Azure Databricks. The job reads data from Azure Data Lake Storage Gen2, performs transformations, and writes results back. You notice that the job is I/O bound. Which THREE strategies can improve performance? (Choose three.)

Hard
12

Your organization uses Azure Data Lake Storage Gen2 for a data lake. You need to prevent accidental deletion of data by enabling a soft delete policy. Which configuration is required?

Medium
13

You store sensitive data in Azure Data Lake Storage Gen2. You need to ensure that only members of a specific security group can read the data, while other users in the organization must not have access, even if they have the Storage Blob Data Reader role at the storage account level. What should you use?

Easy
14

You manage an Azure Data Lake Storage Gen2 account containing a large volume of JSON files. Users report that direct read operations from the data lake are slow, and you observe high egress costs. You need to optimize read performance and reduce cost for analytical queries that frequently filter on a specific timestamp column and select a subset of columns. What should you do?

Medium
15

You are tasked with transforming data in an Azure Synapse Analytics pipeline using a mapping data flow. The source data contains a column 'FullName' in the format 'LastName, FirstName'. You need to split this into two separate columns: 'LastName' and 'FirstName'. Which transformation should you use?

Easy
16

Your team uses Azure Databricks for data processing. You need to implement a cost-control strategy that automatically terminates idle clusters after 30 minutes of inactivity, but allows users to override this policy for specific workloads that require long-running clusters. What is the most efficient approach?

Hard
17

You are migrating a large on-premises SQL Server database to Azure Synapse Analytics. The database includes tables with up to 500 million rows and frequent updates. You need to minimize data movement during the migration while ensuring optimal query performance in the dedicated SQL pool. Which table design strategy should you use?

Hard
18

You have an Azure Databricks workspace that processes sensitive data. The security team requires that all access to the workspace be authenticated using Microsoft Entra ID and that all API calls be audited. Which configuration should you implement?

Medium
19

Your organization uses Azure Data Factory to orchestrate data pipelines. You need to ensure that sensitive data is not exposed in pipeline logs. What should you configure?

Easy
20

You are designing a data processing solution in Azure Synapse Analytics. The solution must process streaming data from Azure Event Hubs and store the results in a dedicated SQL pool. You need to choose the most appropriate service for near real-time ingestion with minimal latency. What should you use?

Medium
21

Your company uses Azure Purview for data governance. You need to ensure that sensitive data in Azure Data Lake Storage Gen2 is automatically detected and classified. What should you configure in Purview?

Medium
22

You are monitoring an Azure Synapse Analytics dedicated SQL pool and notice that queries against a large fact table are slow. The table is distributed using hash distribution on a column that has a high number of nulls. You need to improve query performance. What should you do?

Hard
23

You need to grant a data analyst read access to a specific folder in an Azure Data Lake Storage Gen2 account. The analyst must not be able to read other folders in the same container. You want to follow the principle of least privilege. What should you use?

Easy
24

You are designing an ETL process in Azure Data Factory. You need to transform data using Mapping Data Flows. Which THREE of the following transformations are available in Mapping Data Flows?

Medium
25

You need to design a near-real-time data processing solution that ingests IoT telemetry data from millions of devices. The data must be aggregated per minute and stored in Azure Cosmos DB for low-latency queries. Which Azure service combination should you use?

Medium
26

You are designing a data processing solution using Azure Databricks. The solution must use Delta Lake for ACID transactions and must optimize storage costs by automatically compacting small files. Which feature should you enable?

Hard
27

You need to monitor the performance of your Azure Synapse Analytics dedicated SQL pool. Which metric should you use to identify queued queries due to concurrency limits?

Easy
28

You are troubleshooting a slow-running pipeline in Azure Data Factory that uses a Copy activity to transfer data from Azure Blob Storage to Azure Synapse Analytics. The pipeline processes about 100 GB of CSV files. The copy performance is poor even though the source and sink are in the same region. What is the most likely cause?

Medium
29

A data engineering team is designing a storage solution for a retail company that receives point-of-sale (POS) transaction data from thousands of stores. The data arrives as JSON files in Azure Data Lake Storage Gen2. The team needs to query the data using Azure Synapse Analytics serverless SQL pool and optimize for performance and cost. The data is partitioned by year, month, and day. They want to minimize the amount of data scanned per query. What should they do?

Medium
30

You are running a Python script in Azure Databricks that reads a CSV file from DBFS. The script runs successfully in an interactive notebook but fails when executed as a job with the error: 'Path does not exist: dbfs:/tmp/data.csv'. What is the most likely cause?

Easy
31

A financial services company needs to store transactional data in Azure Cosmos DB. The data is accessed by multiple applications using different partition keys. The company requires strong consistency for financial transactions and wants to minimize latency for reads and writes. Which consistency level should they choose?

Easy
32

You have an Azure Databricks notebook that processes a large Delta table. The notebook uses a structured streaming query to read from the Delta table and write to another Delta table. The source table receives frequent updates and deletes. You need the streaming query to process both new data and changes (updates and deletes) from the source table. What should you do?

Hard
33

Which TWO Azure services can be used to monitor and analyze query performance in Azure Synapse Analytics dedicated SQL pool?

Medium
34

You are designing a data lake architecture using Azure Data Lake Storage Gen2. You need to implement a least-privilege security model. Which authorization mechanism should you use for granular control?

Medium
35

You are designing a data processing solution for a retail company that uses Azure Synapse Analytics. The solution must process point-of-sale (POS) data from multiple stores. The data arrives in CSV files in Azure Data Lake Storage Gen2. Each store sends a file every hour. You need to process the files as they arrive and load the data into a dedicated SQL pool. The solution must handle late-arriving files (files that arrive after the scheduled processing time) and ensure that the data is consistent. Which approach should you use?

Hard
36

You are designing a storage layer for a fraud detection system. The system writes millions of small JSON records per hour to Azure Data Lake Storage Gen2 and must support both batch analytics and interactive queries from Azure Databricks. You need to choose a storage format and layout that minimizes query latency for selective filters on customer ID while keeping storage costs predictable. What should you do?

Hard
37

Refer to the exhibit. A Stream Analytics job shows increasing watermark delay and input deserialization errors. Which action should be taken first to troubleshoot?

Hard
38

Your organization uses Azure Synapse Analytics serverless SQL pool to query Parquet files in Azure Data Lake Storage Gen2. You notice that queries are slow when filtering on a date column. You need to improve query performance without increasing costs. What should you do?

Hard
39

You are developing an Azure Databricks notebook that processes streaming data from Azure Event Hubs using Structured Streaming. The stream writes to a Delta Lake table. You need to ensure that the stream can recover from failures and continue processing from where it left off without reprocessing all data. You also need to minimize the impact on the source. What should you configure?

Medium
40

Your team is migrating an on-premises SQL Server data warehouse to Azure Synapse Analytics. The source has a fact table with 500 million rows and several dimension tables. You need to choose the best distribution strategy for the fact table to minimize data movement during joins. Which distribution type should you use?

Medium
41

Refer to the exhibit. You are deploying an Azure Synapse Analytics dedicated SQL pool using the provided ARM template snippet. After deployment, you need to adjust the performance level to DW200c to handle increased workload. Which parameter should you modify?

Hard
42

Which THREE of the following are best practices for designing tables in a dedicated SQL pool in Azure Synapse Analytics?

Hard
43

Which THREE best practices should be followed when designing a data lake in Azure Data Lake Storage Gen2 for optimal performance?

Easy
44

A data engineering team is designing a batch processing solution using Azure Databricks. The data is stored in Azure Data Lake Storage Gen2 (ADLS Gen2) and must be processed daily with minimal cost. The team needs to choose between using a Delta Lake table or a Parquet file format for the processed output. Which TWO factors should the team consider when making this decision?

Medium
45

You are configuring security for an Azure Data Lake Storage Gen2 account. You need to ensure that users can only access files and folders for which they have explicit permissions, and that permissions are enforced at the file and folder level. What should you enable?

Easy
46

You are running an Azure Stream Analytics job that reads from an Event Hub and writes to a Power BI dataset. The job is falling behind and processing latency is increasing. What should you do to improve performance?

Easy
47

A data engineer needs to store JSON documents that are frequently updated by multiple users concurrently. The solution must support optimistic concurrency control and have built-in indexing on all fields. Which Azure data store should be used?

Medium
48

You need to ensure that data in an Azure Data Lake Storage Gen2 account is encrypted at rest using a customer-managed key. Which feature should you configure?

Easy
49

Your organization uses Azure Data Lake Storage Gen2 and needs to prevent accidental deletion of data by enabling soft delete. You also need to ensure that deleted blobs are recoverable for 30 days. What should you configure?

Easy
50

You are monitoring an Azure Stream Analytics job that processes data from an IoT hub. The job's output to Azure Synapse Analytics is experiencing high latency. The job's SU% utilization is at 90%. Which action will most likely reduce the latency?

Hard
51

Drag and drop the steps to convert data from CSV to Parquet format using Azure Data Factory into the correct order.

Medium
52

You are building a data pipeline that uses Azure Data Factory to copy data from a REST API to Azure Blob Storage. The REST API returns JSON data in pages of 1000 records each. The total number of records is 50,000. Which activity or feature should you use to loop through the pages?

Medium
53

You are designing a data storage solution for a marketing analytics platform. The platform collects clickstream data from websites and needs to store it for both real-time dashboards and historical analysis. The data is semi-structured (JSON) and arrives at a rate of 10,000 events per second. You need to choose an Azure storage solution that can handle the ingestion rate, support schema-on-read, and integrate with Azure Databricks for advanced analytics. The solution must also be cost-effective for long-term storage. What should you use?

Easy
54

Refer to the exhibit. You are deploying an Azure Synapse Analytics workspace using an ARM template. The template defines a managed virtual network integration runtime. You need to ensure that the integration runtime can run mapping data flows with a time-to-live (TTL) of 10 minutes. What is the purpose of the 'timeToLive' property in this configuration?

Medium
55

Which THREE security features are available in Azure Data Lake Storage Gen2 to protect data at rest and in transit? (Choose three.)

Hard
56

You are developing an Azure Synapse Analytics serverless SQL pool solution that queries Parquet files in Azure Data Lake Storage Gen2. Analysts run ad-hoc queries with predicates on a high-cardinality column named TransactionId, and each query scans the entire folder, causing high cost. You need to reduce the amount of data scanned per query without changing the file format. What should you do?

Hard
57

A company uses Azure Synapse Analytics dedicated SQL pool for a data warehouse. They notice that some queries are using more memory than expected, causing resource contention. Which TWO actions should they take to diagnose and optimize memory usage?

Medium
58

Which TWO actions can you take to optimize the performance of a dedicated SQL pool in Azure Synapse Analytics when loading large volumes of data?

Medium
59

Refer to the exhibit. You have a mapping data flow in Azure Data Factory that aggregates sales data. The data flow runs successfully but the sink table contains only the total sum per run instead of per product. What is missing?

Easy
60

A company uses Azure Synapse Analytics dedicated SQL pool. They notice that some queries are slow due to high data movement. What should you do to minimize data movement for queries that join large fact tables?

Medium
61

Which THREE metrics should you monitor for an Azure Synapse Analytics dedicated SQL pool to ensure optimal performance?

Hard
62

You are designing a storage solution for a healthcare analytics platform. The platform ingests large volumes of structured patient records stored as Parquet files in Azure Data Lake Storage Gen2. Analysts query this data using Azure Synapse Analytics serverless SQL pools. To minimize query cost and improve performance, you need to choose an appropriate file organization and table type. What should you do?

Medium
63

You are designing a data storage solution for a media company that ingests video files from various sources into Azure Data Lake Storage Gen2. The files are uploaded continuously and must be processed by Azure Databricks. You need to ensure that the data is organized efficiently for query performance and that access is secure. Which two actions should you include in your design? (Choose two.)

Medium
64

Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to implement a security strategy that allows users to read only specific folders within a container. Which authorization method should you use?

Hard
65

You have an Azure Synapse Analytics dedicated SQL pool that contains a large fact table named FactSales. The table is partitioned by date and has a clustered columnstore index. You notice that queries filtering on a specific date range are slow. You need to improve query performance for these queries. What should you do?

Medium
66

Drag and drop the steps to configure Azure Stream Analytics job with event input and Power BI output into the correct order.

Medium
67

You are building a batch processing solution in Azure Synapse Analytics that reads data from a dedicated SQL pool, applies complex transformations using Synapse Spark, and writes the results back to the dedicated SQL pool. The pipeline must run on a schedule and handle transient failures with retries. Which approach should you use?

Hard
68

You are partitioning a large fact table in Azure Synapse Dedicated SQL Pool by date. The table is used for queries that filter on CustomerID and Date. You want to minimize data movement. Which distribution strategy should you use?

Medium
69

You have an Azure Synapse Analytics workspace with a dedicated SQL pool. You need to create an external table that references Parquet files stored in Azure Data Lake Storage Gen2. The external table will be used for ad-hoc queries. Which statement correctly describes the required components?

Medium
70

You need to transform JSON data containing nested arrays into a tabular format for analysis in Azure Synapse Analytics. Which transformation in Azure Data Factory or Synapse Pipelines should you use?

Easy
71

You need to monitor the performance of an Azure Stream Analytics job in real time. Which Azure service should you use to track the job's resource utilization (e.g., SU % utilization) and set up alerts when the job is approaching its capacity?

Easy
72

You have an Azure Synapse Analytics dedicated SQL pool with a table that uses hash distribution on CustomerID. You notice that queries joining this table with another table on OrderDate are slow. What is the most likely cause?

Hard
73

Which TWO Azure services can be used to monitor Azure Data Factory pipeline runs and set up alerts?

Medium
74

Drag and drop the steps to set up Azure Data Factory pipeline with parameterization and dynamic expressions into the correct order.

Medium
75

Refer to the exhibit. You are reviewing an Azure Stream Analytics job query. The job has a stream input and a reference data input. The job is failing with the error 'Reference data input must be of type Reference, not Stream'. What is the cause of the error?

Easy
76

You are designing a data processing solution in Azure using Azure Data Lake Storage Gen2 as the storage layer. You need to ensure that data ingested from various sources is immutable and can be used for both batch and streaming workloads. Which storage design pattern should you implement?

Medium
77

You are a data engineer for a multinational e-commerce company. The company uses Azure Synapse Analytics as its data warehouse. The current fact table, SalesFact, is distributed using hash distribution on the CustomerID column. It has 2 billion rows and is 2 TB in size. Recently, the business team has been running many queries that aggregate sales by product category and date, and these queries are experiencing high data movement and long execution times. The product dimension table (ProductDim) has 100,000 rows and is 100 MB. The date dimension table (DateDim) has 5,000 rows and is 5 MB. You need to redesign the storage to minimize data movement for these aggregation queries. You cannot change the fact table distribution key to ProductID because of other critical queries that rely on CustomerID. What should you do?

Hard
78

You are developing an Azure Databricks notebook that processes a large Delta Lake table. You must add a derived column that depends on the latest value of a watermark stored in a small reference table, and the notebook must refresh this value before each micro-batch. You need to ensure the reference data is re-read on every micro-batch rather than cached once. Which approach should you use?

Medium
79

Your organization uses Azure Data Lake Storage Gen2 (ADLS Gen2) and wants to transform data using Azure Databricks. The data is stored in Parquet format. You need to read the data into a Spark DataFrame. Which DataFrame reader method should you use?

Easy
80

You have an Azure Synapse Analytics serverless SQL pool. You need to monitor the number of queries that are currently executing. Which dynamic management view should you query?

Easy
81

You have an Azure Data Lake Storage Gen2 account used by an Azure Synapse Analytics serverless SQL pool. Analysts run ad-hoc queries against CSV and Parquet files. You need to reduce the amount of data scanned by these queries without changing file contents. What should you do?

Medium
82

A data engineer needs to store semi-structured JSON log files from a web application. Each log entry is about 1 KB. The logs are rarely queried (once a month) and must be retained for 7 years for compliance. The solution must minimize storage cost. Which storage option should be used?

Easy
83

You are creating an Azure Data Factory pipeline that must copy data from an on-premises SQL Server to Azure Blob Storage daily. The on-premises network restricts inbound connections, and you need a secure connection without exposing the SQL Server to the public internet. What should you use to connect to the on-premises SQL Server?

Easy
84

You need to process a large dataset that contains personally identifiable information (PII). The data must be anonymized before being used for analytics. Which Azure service should you use to apply column-level masking dynamically?

Easy
85

You are a data engineer for a financial services company. The company uses Azure Data Lake Storage Gen2 as its data lake. You have a directory structure where each customer has a folder containing transaction files in CSV format. The security team requires that each customer's data be accessible only to that customer's users. You need to implement fine-grained access control using Azure Data Lake Storage Gen2's POSIX-like ACLs. However, you have thousands of customers, and managing ACLs individually is not feasible. What should you do?

Medium
86

You have an Azure Data Lake Storage Gen2 account that stores large volumes of parquet files. A reporting application frequently queries a specific subset of data filtered by a 'region' column. To minimize query latency and cost, which optimization should you implement?

Medium
87

You are using Azure Synapse Analytics to process data in a dedicated SQL pool. You need to ensure that queries against a large fact table perform well. The fact table is partitioned by date and distributed by a product key. Which two actions should you take? (Choose two.)

Medium
88

You are implementing a data lake using Azure Data Lake Storage Gen2. Which THREE actions should you take to secure the data at rest and in transit?

Medium
89

You are developing a data processing pipeline for a gaming company that uses Azure Databricks. The pipeline processes game event data from Azure Event Hubs. You need to detect cheating patterns by analyzing events in real time. The solution must be able to handle high throughput and low latency. The output should be written to Azure Cosmos DB for real-time dashboards. Which approach should you use?

Medium
90

Which TWO Azure services can be used to audit data access and changes in Azure Data Lake Storage Gen2? (Choose two.)

Easy
91

You are developing an Azure Databricks notebook to process streaming data from Azure Event Hubs. The notebook must write the processed data to a Delta table with exactly-once processing guarantees. You need to configure the write operation. Which option should you use?

Easy
92

You are designing a data processing solution using Azure Databricks. You need to read data from Azure Data Lake Storage Gen2, transform it using Spark SQL, and write to a Delta table. Which TWO configurations are required to ensure optimal performance for large datasets?

Medium
93

A data engineer needs to store semi-structured JSON logs for analysis using Azure Synapse Serverless SQL. Which file format should be used for optimal query performance?

Medium
94

Which TWO Azure services can be used to monitor data pipeline runs and set up alerts for failures in Azure Data Factory?

Hard
95

You need to process streaming data from Azure Event Hubs and store the results in Azure Cosmos DB for a real-time dashboard. The solution must handle duplicate events and ensure exactly-once processing. Which Azure service should you use?

Easy
96

A data engineer monitors an Azure Stream Analytics job that processes real-time data. The job is falling behind, and the SU utilization is at 100%. Which action should be taken to improve performance?

Easy
97

You are designing a data processing solution for a retail company that uses Azure Databricks. The solution needs to process streaming sales data from Event Hubs and batch data from Azure Data Lake Storage Gen2. You need to ensure that the solution can handle late-arriving data and maintain exactly-once semantics. Which TWO technologies should you use?

Hard
98

Which TWO Azure features can be used to encrypt data at rest in Azure Blob Storage? (Choose two.)

Easy
99

You have an Azure Synapse Analytics dedicated SQL pool. A nightly ELT process loads a 500 GB staging table and then applies transformations using a stored procedure. The procedure performs many single-row updates against a large fact table, and the load now exceeds its window. You need to reduce the duration of the transformation step. What should you do?

Hard
100

Your company uses Azure Cosmos DB for NoSQL to store user profiles. The application frequently reads profiles by user ID (the partition key). Occasionally, the application needs to query by email address, which is not part of the partition key. What should you do to optimize the occasional queries by email?

Easy
101

You need to ensure that an Azure Data Factory pipeline retries a failed activity up to three times with a 5-minute delay between retries. How should you configure the activity?

Easy
102

Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to implement a monitoring strategy to detect and alert on unusual access patterns that could indicate a security breach. Which THREE services or features should you use? (Choose three.)

Hard
103

You use Azure Data Lake Storage Gen2 with a hierarchical namespace. You need to delegate permissions to a group of data scientists so they can create folders and upload files only within a specific directory path. What is the best way to achieve this?

Easy
104

You are examining a T-SQL script that creates an external table in Azure Synapse serverless SQL pool. The query SELECT * FROM dbo.Sales returns zero rows, but the folder /year=2024/ in ADLS Gen2 contains Parquet files. What is the most likely cause?

Hard
105

You are configuring security for an Azure Data Lake Storage Gen2 account that stores sensitive data. You need to ensure that all data access is logged and that you can audit who accessed the data and when. You also need to retain the logs for 90 days. What should you do?

Easy
106

You maintain an Azure Stream Analytics job that reads from an Event Hubs input and writes to an Azure Synapse Analytics dedicated SQL pool. During peak load the job produces late-arriving events that are dropped, and downstream reports show missing rows. You must retain and process events that arrive after the watermark by up to several minutes without changing the input. What should you configure?

Hard
107

A financial services company needs to store transaction data for audit purposes. The data must be immutable and cannot be modified or deleted for 7 years. Which Azure storage feature should be used?

Hard
108

You are designing a batch processing solution for a data lake. Source files arrive daily in Parquet format in Azure Data Lake Storage Gen2. The data must be cleaned, aggregated, and loaded into an Azure Synapse SQL pool. The solution should minimize compute costs and management overhead. Which technology should you use for the transformation?

Easy
109

You are a data engineer for a healthcare company that processes patient data. You have an Azure Databricks workspace with a cluster configured for data processing. You need to implement a solution that processes streaming data from Azure Event Hubs, enriches it with reference data stored in Azure Cosmos DB, and writes the output to Delta Lake in Azure Data Lake Storage Gen2. The solution must ensure that the data processing is fault-tolerant and can handle schema evolution. The reference data is updated infrequently. You need to choose an approach that minimizes complexity and cost. What should you do?

Hard
110

A financial services firm stores trade records in an Azure Data Lake Storage Gen2 account. Regulatory requirements mandate that all data at rest be encrypted with a customer-managed key (CMK) stored in Azure Key Vault, and that the key be rotated every 90 days without re-uploading any data. The storage account currently uses Microsoft-managed keys. What should you do to meet these requirements with the least administrative effort?

Medium
111

Your company uses Azure Data Lake Storage Gen2. You need to ensure that data at rest is encrypted using a customer-managed key stored in Azure Key Vault. What should you configure?

Easy
112

Refer to the exhibit. A data engineer wants to copy only new orders from an Azure SQL database to Azure Data Lake Storage Gen2. The pipeline runs daily at midnight. What should be added to the pipeline to ensure incremental loads?

Medium
113

You manage an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline runs daily and completes successfully. You need to be alerted when the pipeline duration exceeds 60 minutes. You want to minimize administrative effort. What should you do?

Medium
114

You are designing a data storage solution for a global IoT application that ingests millions of events per second. The data is write-heavy with occasional reads for real-time dashboards. Which Azure storage option and configuration would provide the lowest latency writes with high throughput?

Hard
115

A company is designing a data storage solution for streaming IoT telemetry data. The data is JSON-formatted, arrives at up to 10,000 events per second, and must be stored for at least 30 days for real-time dashboards and ad-hoc querying. The solution must minimize operational overhead and query latency. Which Azure service should they use?

Medium
116

You are responsible for securing an Azure Synapse Analytics workspace that contains sensitive data. You need to ensure that data is encrypted at rest using a customer-managed key stored in Azure Key Vault. What should you configure?

Easy
117

You need to store JSON files from an external partner in Azure Blob Storage. The files contain sensitive financial data. Which access method provides the highest security while allowing the partner to upload files?

Easy
118

You are designing a data lake in Azure Data Lake Storage Gen2 for a large enterprise. You need to ensure that only authorized users can access the data, and you must implement the principle of least privilege. Which security mechanism should you use to grant fine-grained access to specific directories and files without modifying the underlying storage account firewall settings?

Hard
119

You are designing a streaming data solution for IoT devices that generate 10,000 events per second. The data must be processed with sub-second latency and then stored in Azure Data Lake Storage Gen2 for archival. Which Azure service should you use for the stream processing?

Medium
120

You need to design a storage solution for IoT device telemetry data that will be queried by time range. The data is append-only and arrives at high velocity. Which TWO features should you use to optimize query performance and reduce costs?

Easy
121

You are designing a data processing solution for an e-commerce company that uses Azure Synapse Analytics. The solution must process clickstream data from a web application. The data arrives in JSON format through Azure Event Hubs. You need to load the data into a dedicated SQL pool every 5 minutes with minimal latency. The data volume is about 100 MB every 5 minutes. You want to use PolyBase for loading. Which approach should you use?

Medium
122

You are optimizing a pipeline in Azure Data Factory that copies data from Azure Blob Storage to Azure Synapse Analytics. The pipeline uses a copy activity with PolyBase. The data is partitioned by date in Blob Storage. You notice that the load is slow. What is the most likely cause?

Hard
123

You are tuning a dedicated SQL pool in Azure Synapse Analytics. A query that joins two large tables (fact_sales and dim_product) is slow. The fact_sales table is hash-distributed on product_id, and dim_product is replicated. You notice that the query plan shows a shuffle move. What is the most likely cause?

Hard
124

You are developing an Azure Databricks notebook that reads a large Delta table, performs a join with a smaller reference table, and writes the result back to Delta Lake. The job runs on a cluster with autoscaling enabled and frequently spills to disk during the join. You need to reduce shuffle and improve performance without changing the result. Which action should you take?

Hard
125

A healthcare company stores patient records in Azure Data Lake Storage Gen2. The data must be organized for efficient querying by a Synapse Analytics serverless SQL pool. The data is currently stored as many small CSV files in a flat directory. You need to improve query performance and reduce the cost of scanning unnecessary data. What should you do?

Easy
126

You have an Azure Data Lake Storage Gen2 account that contains sensitive data. You need to ensure that data is encrypted at rest and that you control the encryption keys. You also need to be able to audit key usage. What should you implement?

Easy
127

Which TWO Azure services can be used to perform real-time data processing on streaming data?

Medium
128

Which TWO techniques can you use to handle schema drift in Azure Data Factory mapping data flows?

Easy
129

Refer to the exhibit. You have an Azure Data Factory pipeline that copies data from a CSV file in Blob Storage to a Synapse dedicated SQL pool table named dbo.Sales. The pipeline fails. The error message indicates that the 'Amount' column in the sink table does not allow NULLs but the source contains NULL values. What is the best way to resolve this issue without losing data?

Hard
130

You are implementing a mapping data flow in Azure Data Factory that joins a large fact table in Azure Synapse Analytics with a slowly changing dimension (SCD) table in Azure SQL Database. The fact table has 500 million rows and the dimension has 2 million rows. You need to optimize the join performance and minimize data movement. The dimension table is small enough to fit in memory. Which join type should you configure in the data flow?

Hard
131

Your Azure Synapse Analytics dedicated SQL pool is experiencing performance degradation. You notice that some queries are being queued due to resource class conflicts. What should you implement to optimize performance and reduce queuing?

Medium
132

You are building an Azure Data Factory pipeline that calls an external REST API returning a JSON array of records. The API paginates results using a 'nextLink' field in the response body, and the number of pages varies per run. You must ingest all pages into Azure Blob Storage in a single pipeline run. Which activity configuration should you use?

Medium
133

You are building an Azure Stream Analytics job that reads from an Azure Event Hub capturing device telemetry. The job must emit results into an Azure Synapse Analytics dedicated SQL pool. You need to minimize latency and avoid intermediate storage. What should you do?

Medium
134

You have an Azure Data Factory pipeline that copies data from an FTP server to Azure Blob Storage. The pipeline runs successfully most of the time, but occasionally fails with a 'FTP server connection refused' error during peak hours. You need to minimize these failures with minimal cost. What should you do?

Easy
135

You are designing a data processing solution using Azure Synapse Analytics serverless SQL pool. The solution will query data stored in Parquet files in Azure Data Lake Storage Gen2. You need to ensure that the queries are optimized for performance. Which action should you take?

Hard
136

You need to store semi-structured JSON data from a web application. The data schema may change over time. The solution must support low-latency queries and be globally distributed. Which Azure data service should you use?

Easy
137

Drag and drop the steps to implement incremental data loading using Azure Data Factory into the correct order.

Medium
138

You are a data engineer at a healthcare company. Your Azure Synapse Analytics workspace contains a dedicated SQL pool that holds patient records. A new compliance rule requires that all queries against the dedicated SQL pool be audited, and that any attempt to access data from an unauthorized IP address be logged. You need to configure auditing for the dedicated SQL pool. What should you do?

Medium
139

You are building an Azure Synapse Analytics pipeline that processes JSON files landing in Azure Data Lake Storage Gen2. The files contain nested arrays representing order line items. You need to flatten this nested structure into a tabular format within a Mapping Data Flow before loading to a dedicated SQL pool. The solution must minimize data movement and avoid writing intermediate files to storage. Which transformation should you use to flatten the nested arrays?

Medium
140

Which THREE options are valid ways to transform data in Azure Synapse Analytics?

Medium
141

You are tasked with designing a data storage solution for a social media analytics company. They need to store user profile data (JSON) and social media posts (text and images). The data is used for machine learning models that require fast random access to individual user profiles and the ability to run analytical queries over posts. The solution must provide low-latency reads for user profiles (milliseconds) and support for large-scale analytics on posts. Which combination of Azure data services should you recommend?

Easy
142

You are designing a data processing solution in Azure Synapse Analytics. The solution must process streaming data from Azure Event Hubs and store the results in a dedicated SQL pool. The solution must support exactly-once semantics and handle late-arriving data. Which Azure service should you use to implement this solution?

Medium
143

Your company is migrating an on-premises SQL Server database to Azure SQL Database. The database includes a large fact table with hourly updates. You need to minimize downtime during migration. Which Azure service should you use to replicate data continuously?

Easy
144

Refer to the exhibit. A data engineer creates an external table in Azure Synapse Analytics pointing to Parquet files in ADLS Gen2. The query 'SELECT * FROM Sales' returns 0 rows, but the files exist. What is the most likely cause?

Medium
145

You are a data engineer at a healthcare company. You have an Azure Data Lake Storage Gen2 account named sthealthcare with a container named records. The container holds sensitive patient data in Parquet files. You need to ensure that only users who are members of the Azure AD group named ClinicalResearchers can read the data, while users in the group DataEngineers can read and write. Access must be managed at the directory level and must not affect other containers in the storage account. What should you do?

Medium
146

A company is designing a data storage solution for IoT device telemetry data. The data is append-only, needs to be stored cost-effectively for long-term analytics, and must support querying by device ID and timestamp. Which Azure storage solution should they use?

Easy
147

A data engineering team is designing a batch processing pipeline that reads from Azure Data Lake Storage Gen2, transforms data using Azure Databricks, and writes to Azure Synapse Analytics. The pipeline must process data incrementally and handle late-arriving data up to 2 hours. Which approach should they use to track processed files?

Medium
148

Your organization uses Azure Purview for data governance. You need to ensure that sensitive data is properly classified and that access to it is monitored. Which THREE actions should you take? (Choose three.)

Hard
149

You are designing a data storage solution in Azure Synapse Analytics dedicated SQL pool. The solution must support efficient loading of large volumes of data from external sources and provide high query performance for reporting. You need to choose two table distribution types that are most appropriate for large fact tables and dimension tables respectively. (Choose two.)

Medium
150

You are building an Azure Stream Analytics job that processes JSON telemetry from Azure Event Hubs. The events contain a nested array field named `readings` with sensor values. You need to transform the data so that each sensor reading becomes a separate output row, and then write the results to Azure Synapse Analytics. Which two actions should you perform? (Choose two.)

Medium
151

You need to monitor the performance of an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline runs on a self-hosted integration runtime. Which metric is most important to monitor to ensure the self-hosted IR is not a bottleneck?

Easy
152

A data engineer needs to store CSV files containing customer data in Azure Blob Storage. The files must be encrypted at rest using a customer-managed key stored in Azure Key Vault. What should they configure?

Easy
153

You are implementing a medallion architecture in Azure Databricks. The silver layer must contain deduplicated, conformed records, and the gold layer must serve aggregated reporting tables. You need to choose Delta Lake operations that support incremental, idempotent updates as new bronze files arrive. Which two operations should you use? (Choose two.)

Hard
154

An organization is using Azure Synapse Analytics and wants to implement column-level security to restrict access to sensitive columns. Which feature should they use?

Easy
155

You are using Azure Data Lake Storage Gen2 as the data lake for your organization. You need to process files in the 'incoming' folder using a scheduled Azure Databricks notebook. After processing, the files should be moved to the 'processed' folder. The files are large (up to 10 GB) and you want to minimize the time to move them. Which approach should you use?

Medium
156

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline uses a self-hosted integration runtime. You need to ensure that data is encrypted in transit and that the integration runtime authenticates to the on-premises SQL Server using Windows authentication. What should you configure?

Hard
157

Drag and drop the steps to set up Azure Purview for data cataloging and lineage tracking into the correct order.

Medium
158

You are implementing a real-time analytics solution using Azure Stream Analytics. The job ingests data from Azure Event Hubs and must output to an Azure SQL Database. You need to ensure that the job can handle out-of-order events and produce accurate aggregations over 5-minute windows. Which setting should you configure?

Hard
159

You are designing a storage solution for a financial analytics platform that ingests CSV files into Azure Data Lake Storage Gen2. Analysts run complex T-SQL queries against the data using Azure Synapse Analytics serverless SQL pool. You need to minimize query cost and improve performance. What should you do?

Medium
160

Your team runs Azure Data Factory pipelines that must copy files from an on-premises file share to Azure Data Lake Storage Gen2 on a nightly schedule. The on-premises network blocks inbound connections and the data must not be exposed to the public internet. You need to enable connectivity without opening firewall ports. What should you deploy?

Easy
161

You manage an Azure Synapse Analytics workspace. A dedicated SQL pool contains a table with a column named CustomerEmail that stores email addresses. You need to ensure that users who are not members of the DataPrivacy role see only a masked version of the email addresses when they query the table, while members of DataPrivacy see the actual values. The solution must minimize administrative effort. What should you do?

Medium
162

You are reviewing a script to create an external data source in Azure Synapse Analytics serverless SQL pool. Based on the exhibit, what is the purpose of the SAS token?

Medium
163

A media company uses Azure Data Lake Storage Gen2 to store video files and metadata. They need to ensure that when a user is deleted from Azure Active Directory, their access to the data lake is immediately revoked. They also want to minimize administrative overhead. What should they do?

Medium
164

You are a data engineer at a healthcare analytics company. The company uses Azure Data Factory (ADF) to orchestrate data pipelines that ingest patient data from on-premises SQL Server databases into Azure Synapse Analytics. Recently, the pipeline has been failing intermittently with the following error: 'Failure happened on 'Sink' side. ErrorCode=SqlFailedToConnect, Type=Microsoft.DataTransfer.Common.Shared.HybridDeliveryException, Message=Cannot connect to SQL Server Database. The TCP connection to the host <server_name>, port 1433 has failed. Error: 'Connection timed out.'.' The on-premises SQL Server is behind a corporate firewall. The ADF self-hosted integration runtime (SHIR) is installed on a VM inside the corporate network. You have verified that the SHIR is running and that the SQL Server is accessible from the SHIR VM using SQL Server Management Studio (SSMS). The error occurs sporadically, not consistently. What is the most likely cause of the intermittent connection timeout?

Medium
165

You are creating an Azure Synapse Analytics pipeline that must copy data from an Azure SQL Database into a dedicated SQL pool. The destination table already exists and the pipeline must append new rows without truncating existing data. Which staging and load option should you configure in the Copy activity?

Easy
166

Which TWO strategies can be used to optimize storage costs for historical data in Azure Data Lake Storage Gen2?

Hard
167

Which THREE are best practices for optimizing query performance in Azure Synapse Analytics dedicated SQL pool?

Hard
168

You are analyzing a Kusto query in Azure Data Explorer that calculates total sales per product for January 2024 and filters for products with sales over 10,000. The query uses the materialize() function. You notice that the query runs slower than expected. What is the primary reason the materialize() function may not be providing the expected performance benefit in this query?

Hard
169

You are optimizing a Spark DataFrame transformation in Azure Synapse Analytics. The DataFrame has 20 columns and 100 million rows. You notice that the job is slow due to many small files being written to the output. Which two actions can you take to reduce the number of output files? (Choose two.)

Easy
170

You are building an Azure Stream Analytics job that reads JSON telemetry from an Azure Event Hub, calculates a 5-minute tumbling window average per device, and writes results to an Azure Synapse Analytics dedicated SQL pool. The stream must handle occasional bursts of late-arriving events by including events that arrive up to 3 minutes after the window closes. You need to configure the job's event ordering settings to meet the late-arrival requirement while minimizing memory usage. What should you do?

Medium
171

You are designing a data transformation solution for a retail company. The company receives daily CSV files from 200 stores via SFTP. The files must be cleaned, validated, and aggregated before loading into Azure Synapse dedicated SQL pool. The solution must minimize administrative overhead and support easy monitoring. Which approach do you recommend?

Medium
172

Which TWO features are available in Azure Data Lake Storage Gen2 but not in Azure Blob Storage? (Choose two.)

Easy
173

You need to monitor the performance of an Azure Stream Analytics job that processes real-time IoT data. Which metric indicates the number of events that are being dropped or delayed due to insufficient processing capacity?

Easy
174

You are designing a near-real-time analytics pipeline for a retail company. Transaction data is generated in Azure SQL Database and must be replicated to Azure Synapse Analytics (dedicated SQL pool) with less than 5 minutes latency. The source table has 50 million rows and 200 columns, but only 30 columns are needed for analytics. Which approach should you recommend?

Hard
175

A manufacturing company uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled and Azure Databricks for analytics. The security team requires that all data stored in the 'raw' container be encrypted at rest using customer-managed keys. The data is ingested via Azure Data Factory. What should the data engineer configure to meet the requirement?

Easy
176

You have an Azure Synapse Analytics dedicated SQL pool. You notice that some queries are taking longer than expected. After reviewing the query plans, you see that some queries are spilling to tempdb. What should you do to reduce tempdb spills?

Medium
177

A data engineer is designing a solution to store historical sales data for a retail company. The data is append-only and accessed infrequently for compliance reports. The solution must minimize storage costs while allowing retrieval within 24 hours. Which storage tier should be used for the data?

Medium
178

A company uses Azure Synapse Analytics dedicated SQL pool to store sales data. The fact table is partitioned by date and distributed by product ID. Queries often join the fact table with a small dimension table on product ID. You notice that these joins cause significant data movement. You need to minimize data movement for these joins. What should you do?

Hard
179

Which THREE metrics should you monitor to evaluate the performance of an Azure Stream Analytics job?

Hard
180

Which TWO of the following are supported sources for Azure Data Factory Copy activity? (Choose two.)

Easy
181

You have a dedicated SQL pool in Azure Synapse that stores a fact table with over 100 billion rows. Query performance is degrading over time. You notice that the table is hash-distributed on a column with many duplicate values. What is the most likely impact?

Medium
182

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline uses a self-hosted integration runtime and runs successfully during business hours. However, after a recent network security update, the pipeline fails with a connection error to the on-premises SQL Server. What is the most likely cause?

Easy
183

You are designing a storage solution for a financial analytics platform. The data consists of large Parquet files stored in Azure Data Lake Storage Gen2. Analysts run complex queries that scan entire partitions, but only a subset of columns is needed for each query. You need to minimize the amount of data read from storage and improve query performance. What should you do?

Medium
184

You are designing a data processing solution in Azure Synapse Analytics. The solution must ensure that sensitive columns containing personally identifiable information (PII) are masked at query time for users without explicit permissions. Which Azure Synapse Analytics feature should you use?

Medium
185

You are designing a security strategy for an Azure Data Lake Storage Gen2 account that stores sensitive data. You need to ensure that data is encrypted at rest using customer-managed keys. What should you configure?

Easy
186

You are designing a storage layer in Azure Data Lake Storage Gen2 for a data engineering pipeline. You need to store Parquet files that will be queried by Azure Synapse Analytics serverless SQL pools and Azure Databricks. You must optimize for query performance and minimize data scanned. Which two actions should you perform? (Choose two.)

Hard
187

You are optimizing a data pipeline in Azure Synapse Analytics that loads data from a CSV file in ADLS Gen2 into a dedicated SQL pool using PolyBase. The load is slow and you need to improve performance. Which action would be MOST effective?

Hard
188

Your team has deployed an Azure Stream Analytics job that writes output to Azure Cosmos DB. You need to monitor the job for data latency and ensure it meets a service-level agreement (SLA) of under 10 seconds from input to output. Which metric should you track in Azure Monitor?

Easy
189

Which of the following are valid activities in an Azure Data Factory pipeline? (Choose three.)

Easy
190

Your organization uses Microsoft Purview to catalog data assets. You need to ensure that sensitive data such as credit card numbers are automatically detected and labeled. Which Purview feature should you configure?

Easy
191

Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to grant a service principal read and write access to a specific directory without granting access to the parent directories. What should you use?

Hard
192

You are implementing a mapping data flow in Azure Data Factory that processes data from an Azure SQL Database. The data flow includes a derived column transformation that adds a new column based on a complex expression. You need to ensure the expression handles null values appropriately. Which function should you use to replace null values with a default?

Medium
193

You are designing a data storage solution in Azure Synapse Analytics. You need to load data incrementally from an Azure Data Lake Storage Gen2 source into a dedicated SQL pool. The source files are appended daily with new data, and you must ensure that only new records are loaded without duplicating existing records. The target table has a column named LoadDate that records when each row was inserted. Which approach should you use?

Hard
194

You are designing a storage solution in Azure Synapse Analytics for a financial services company. The company ingests trade data into a dedicated SQL pool. The data is partitioned by trade date and queried primarily by date ranges. To improve query performance and reduce data movement, you need to choose an appropriate distribution type for the fact table. The table is large (over 2 billion rows) and frequently joined with a smaller dimension table on a non-distributed key. Which distribution type should you use?

Medium
195

Your team uses Azure Synapse Analytics serverless SQL pool to query Parquet files in Azure Data Lake Storage Gen2. The query performance is inconsistent, and some queries take a long time to execute. You need to improve query performance. What should you do?

Hard
196

A data engineering team is building a batch processing solution for a financial services company. Data is ingested daily from multiple sources into Azure Data Lake Storage Gen2 in CSV format. The data must be transformed (filtered, aggregated, joined) and loaded into Azure Synapse Analytics dedicated SQL pool. The team must optimize for cost and performance. The total data volume is 2 TB per day. The team has the following options: Option A: Use Azure Data Factory pipelines with copy activity to load raw CSV files into Synapse staging tables, then use T-SQL stored procedures in Synapse to perform transformations. Option B: Use Azure Databricks with Auto Loader to incrementally ingest CSV files, perform transformations in Spark, and write the results to Synapse using the Spark Synapse connector. Option C: Use Azure Data Factory with mapping data flows to transform the data in a serverless environment and then write to Synapse. Option D: Use Azure Synapse Pipelines (built on ADF) with a notebook activity that runs a PySpark notebook in Synapse Spark pool to transform and load data. Which option should the team choose to minimize cost and management overhead while meeting performance requirements?

Medium
197

A data engineer needs to process a large dataset stored in Azure Blob Storage using Azure Databricks. The dataset consists of millions of small CSV files. The processing job is slow due to the overhead of reading many small files. Which technique should be used to improve performance?

Easy
198

You are monitoring an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline runs daily and has recently started taking longer than expected. You need to identify the cause of the performance degradation. Which two actions should you perform? (Choose two.)

Medium
199

You are monitoring an Azure Data Factory pipeline that runs hourly. The pipeline executes a stored procedure in an Azure SQL Database. Recently, you have observed that the pipeline occasionally fails with a 'Deadlock' error when the stored procedure runs. The Azure SQL Database is configured with the 'Read Committed Snapshot' isolation level enabled. You need to resolve the deadlock issue with minimal impact on performance. The stored procedure updates multiple tables in a single transaction and is critical for reporting. What should you do?

Easy
200

You are implementing dynamic data masking on an Azure Synapse Analytics dedicated SQL pool. A table named Customers contains columns: CustomerID (int), Email (varchar), Phone (varchar), and CreditCard (varchar). You need to mask the Email and Phone columns so that users without elevated permissions see only the last four characters of the Email and Phone, while users with elevated permissions see the full values. You also need to ensure that the masking does not affect the storage size of the columns. What should you do?

Hard
201

You are designing a real-time analytics solution for IoT devices that emit telemetry data every second. The data must be aggregated every minute and stored in Azure SQL Database for historical analysis. You need to minimize latency and operational overhead. Which approach should you recommend?

Hard
202

You are an administrator for an Azure Synapse Analytics dedicated SQL pool. You execute the T-SQL statements shown in the exhibit. The external table 'dbo.Orders' is created. Which statement about querying this external table is true?

Easy
203

Your organization is using Azure Synapse Analytics dedicated SQL pool. You notice that queries are running slower than expected. Upon reviewing the execution plans, you see that some queries are performing table scans instead of seeks on large fact tables. What is the most likely cause?

Medium
204

Which TWO methods can you use to optimize the cost of storing data in Azure Data Lake Storage Gen2?

Easy
205

You are developing an Azure Data Factory pipeline that must call an external REST API, parse the JSON response, and load selected fields into an Azure SQL Database. The API requires a bearer token that expires every hour, so the pipeline must obtain a fresh token before each call. You need to implement the token acquisition and header injection without writing custom code in a data flow. What should you use?

Easy
206

You are designing a data storage solution for a company that uses Azure Data Lake Storage Gen2. The company needs to store data in a hierarchical namespace and requires fine-grained access control at the folder and file level. The solution must also support POSIX-style permissions. Which feature should you enable?

Easy
207

A company uses Azure Databricks for data processing. They want to monitor the performance of Spark jobs and set up alerts for job failures. Which Azure service should they use?

Medium
208

You are designing a storage solution for a healthcare analytics platform. The data lands in Azure Data Lake Storage Gen2, and analysts must query it through Azure Synapse Analytics serverless SQL pools. Security policy requires that analysts see only the columns relevant to their role, and that access be governed by Microsoft Entra ID identities rather than shared keys. Which two actions should you include in the design? (Choose two.)

Medium
209

You are a data engineer at a financial services company. The company uses Azure Cosmos DB for NoSQL to store customer transaction data. The data is partitioned by customerId. The application team needs to run analytical queries that aggregate transactions by date across all customers. These queries are currently slow and consume high RUs. You need to enable faster analytical queries without impacting the transactional workload. What should you do?

Easy
210

You are designing a data storage solution for real-time analytics on IoT telemetry. The system must ingest 10,000 events per second and support sub-second query latency. Which Azure data store should you use?

Medium
211

You are implementing a Spark Structured Streaming job in Azure Databricks that reads from an Azure Event Hubs topic. The job must handle late-arriving data up to 10 minutes and produce aggregated results every 5 minutes. You need to configure the watermark and window. Which code snippet should you use?

Hard
212

You need to monitor the health of your Azure Data Lake Storage Gen2 account. Which metric should you use to track the number of successful and failed requests?

Easy
213

You manage an Azure Data Lake Storage Gen2 account containing a large volume of JSON logs. Users frequently query only the last seven days of data, but each query scans the entire dataset, causing high costs and slow response times. You need to reduce the amount of data scanned by queries without changing the data format or moving the data. What should you do?

Medium
214

You are designing a data processing solution in Azure that must handle both batch and streaming data. The solution should use a common storage layer for both and support schema evolution. Which TWO technologies should you recommend?

Medium
215

Your organization needs to ensure that all data stored in Azure Data Lake Storage Gen2 is encrypted at rest using Microsoft-managed keys. What is the default encryption method?

Easy
216

Your team is developing a data processing solution in Azure Synapse Analytics. You need to ensure that the solution can automatically scale compute resources based on workload demand for serverless SQL pools. Which feature should you configure?

Easy
217

You need to incrementally load new and updated records from a source SQL Server database to Azure Synapse Dedicated SQL Pool. The source table has a LastModifiedDate column. Which Azure Data Factory feature should you use to implement incremental loading efficiently?

Easy
218

You are designing a data pipeline in Azure Synapse Analytics to ingest data from Azure Blob Storage into a dedicated SQL pool. The source files are CSV with varying row lengths, and you need to ensure optimal performance for reads. Which file format and compression should you recommend?

Medium
219

You are creating an Azure Data Factory data flow that transforms data from an Azure SQL Database. The data flow must filter rows based on a column value and then aggregate the results. Which transformation should you use first?

Easy
220

You are developing a data processing solution that requires aggregating sales data from multiple CSV files stored in Azure Data Lake Storage Gen2. The data should be cleansed and transformed before loading into Azure Synapse Analytics. Which Azure service should you use to implement a code-free transformation pipeline?

Easy
221

You are designing a data storage solution for a retail company that expects high volumes of small, time-series sensor data from thousands of IoT devices. The data must be stored cost-effectively and queried by time range with low latency. Which Azure data store should you recommend?

Medium
222

You are implementing a solution in Azure Databricks that reads from a Delta Lake table and writes to another Delta Lake table. The pipeline must process data incrementally and handle updates and deletes from the source. Which feature should you use to read the changes?

Hard
223

Your organization uses Azure Synapse Analytics dedicated SQL pool. You need to ensure that all data at rest in the SQL pool is encrypted using a customer-managed key stored in Azure Key Vault. What should you configure?

Medium
224

You have an Azure Data Lake Storage Gen2 account that contains sensitive data. You need to implement a solution that enforces access control at the file and folder level, and also allows you to audit access. You want to minimize administrative effort. What should you do?

Hard
225

You are designing a data storage solution in Azure Synapse Analytics. You need to store large fact tables that are frequently joined with dimension tables on a common column. The solution must minimize data movement during query execution and support high-concurrency queries. Which two actions should you take? (Choose two.)

Medium
226

You have an Azure Synapse Analytics dedicated SQL pool that handles both high-priority real-time queries and low-priority batch jobs. You need to ensure that high-priority queries always get the resources they need, while batch jobs do not starve. What should you configure?

Medium
227

You are implementing a Mapping Data Flow in Azure Synapse Analytics that reads from a Parquet source and writes to a Delta sink. The data flow includes a Surrogate Key transformation to generate unique keys for each row. You notice that when the data flow runs multiple times, the surrogate keys are not consistent across runs and sometimes overlap. You need to ensure that surrogate keys are unique and stable across runs. What should you do?

Hard
228

Which TWO of the following are required components to set up a data pipeline that uses Change Data Capture (CDC) to incrementally load data from SQL Server to Azure Synapse using Azure Data Factory?

Easy
229

You are building a real-time dashboard to monitor user activity on a website. The data is ingested via Azure Event Hubs and must be aggregated every minute with a 30-second late-arrival tolerance. The aggregated results should be stored in Azure Cosmos DB for low-latency reads. Which Azure service should you use to perform the windowed aggregation?

Medium
230

You are a data engineer at a large retail company. Your team uses an Azure Synapse Analytics workspace with a dedicated SQL pool. You need to implement row-level security (RLS) so that sales representatives can see only data for their own region. You must ensure that the security predicate is evaluated at query time and that users cannot bypass it by using different tools. What should you do?

Medium
231

You are designing a data processing solution in Azure Data Factory that uses mapping data flows. You need to perform type conversions on incoming data. Which two transformations can be used to change data types? (Choose two.)

Easy
232

Your team is using Azure Synapse Analytics to process sensitive customer data. You need to ensure that column-level security is applied to a specific table so that only users with a certain role can view certain columns. Which feature should you use?

Medium
233

You are designing a data processing solution for a financial services company. The solution must process sensitive customer data and comply with GDPR. The data will be stored in Azure Synapse Analytics. You need to ensure that only authorized users can view specific columns (e.g., credit card numbers). Which security feature should you implement?

Hard
234

You are building an Azure Stream Analytics job that reads from an Azure Event Hubs input and writes to an Azure Synapse Analytics dedicated SQL pool. You need to compute a 5-minute tumbling window aggregation that outputs only once per window after all events for that window have arrived. Which query construct should you use?

Medium
235

You are developing an Azure Databricks notebook that processes streaming data from Azure Event Hubs and writes to a Delta Lake table. The stream must handle late-arriving data up to 30 minutes old and ensure that aggregations are computed correctly even if events arrive out of order. You need to minimize state store size and avoid unbounded growth. Which combination of features should you use?

Hard
236

You are designing a data processing solution in Azure Synapse Analytics. The solution must use a dedicated SQL pool to store fact and dimension tables. The fact table is expected to have billions of rows. Which distribution strategy should you recommend for the fact table to optimize query performance and minimize data movement?

Easy
237

You are optimizing an Azure Synapse Analytics dedicated SQL pool that stores a large fact table. Queries frequently join the fact table to a small dimension table on a non-distributed column, causing data movement. You need to reduce data movement and improve query performance. (Choose two.)

Medium
238

Refer to the exhibit. An ARM template deploys an Azure Synapse Analytics workspace. What is the purpose of the 'managedVirtualNetwork' property set to 'default'?

Medium
239

You are a data engineer at a logistics company. You have an Azure Data Lake Storage Gen2 account that stores JSON logs from IoT devices. The logs are written continuously and are stored in a folder structure of /logs/{year}/{month}/{day}/{hour}/. You need to optimize the storage for cost and performance. The data is accessed frequently for the first 30 days, then occasionally for the next 60 days, and rarely after that. You need to minimize storage costs while ensuring that data remains available. What should you do?

Medium
240

Refer to the exhibit. You are creating a serverless SQL table in Azure Synapse Analytics that reads Parquet files from the specified location. The folder contains multiple Parquet files with different schemas. When querying the table, you get an error about schema mismatch. What is the most likely reason?

Hard
241

You are designing a data pipeline in Azure Data Factory that copies data from Azure Blob Storage to Azure SQL Database. The data contains personally identifiable information (PII). What should you use to protect the data during transit?

Easy
242

You are implementing a Spark Structured Streaming job in Azure Databricks that consumes from an Azure Event Hubs topic and writes to a Delta table. The stream must tolerate reprocessing after a cluster restart without producing duplicate rows in the Delta table. You need to configure the write path accordingly. (Choose two.)

Hard
243

You are designing a data pipeline that uses Azure Data Factory to load data from an FTP server to Azure Data Lake Storage. The FTP server requires authentication with username and password. Which type of linked service should you create?

Easy
244

Your company uses Azure Databricks for data processing. You need to ensure that spark jobs cannot access certain storage accounts. What is the most secure approach?

Easy
245

You are creating an Azure Data Factory pipeline that must copy data from an on-premises Oracle database to Azure Blob Storage every night. The on-premises network restricts inbound connections. You need to configure the integration runtime. What should you do?

Easy
246

You are authoring an Azure Databricks notebook that reads Parquet files from Azure Data Lake Storage Gen2 and must write results to a Delta table. Users report that queries against the Delta table return stale data after each notebook run, even though the write succeeds. You need to ensure readers always see the latest committed data. What should you do?

Medium
247

Which TWO of the following are valid methods to secure data at rest in Azure Data Lake Storage Gen2?

Medium
248

You are running a pipeline in Azure Data Factory that uses a Mapping Data Flow. The data flow reads from Azure SQL Database and writes to Azure Synapse Analytics. You find that the data flow is very slow. Which configuration change would most likely improve performance?

Medium
249

You are designing a data storage solution for a healthcare analytics platform. The solution must store patient records in Azure SQL Database and allow point-in-time restore for any time within the last 35 days. The data must be encrypted at rest using customer-managed keys (CMK) stored in Azure Key Vault. You need to configure the Azure SQL Database to meet these requirements. What should you do?

Medium
250

You are using Azure Databricks to process a large dataset stored in Delta Lake. You need to reduce the number of files scanned during queries by organizing data into folders based on a commonly filtered column. Which Delta Lake feature should you implement?

Easy
251

A multinational bank needs to store customer transaction records for 10 years to meet regulatory compliance. The data is rarely accessed after the first year. The solution must minimize storage costs while allowing queries on recent data with low latency. Which tiering strategy should you implement?

Hard
252

You are responsible for securing an Azure Synapse Analytics workspace. You need to ensure that only authorized users can query the serverless SQL pool. Which authentication method should you use?

Medium
253

You are designing a data processing solution that uses Azure Databricks to transform large datasets. You need to ensure that the processing is cost-effective and can scale to handle variable workloads. Which cluster configuration should you recommend?

Medium
254

A company wants to ingest streaming data from IoT devices into Azure for real-time analytics. The data must be available for immediate querying and also stored long-term in a cost-effective format. Which Azure service should be used as the primary ingestion endpoint?

Easy
255

You are monitoring an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Synapse Analytics using a self-hosted integration runtime. You notice that the pipeline runs are taking longer than expected, and you suspect performance bottlenecks. You need to identify the cause and optimize the copy performance. What should you do first?

Hard
256

You are designing a data processing solution for a marketing company that uses Azure Synapse Analytics. The solution needs to process customer data from multiple sources, including CRM and web analytics. The data must be cleansed and transformed before loading into a dedicated SQL pool. The transformations include string manipulations, date conversions, and lookups. You need to choose a serverless transformation approach that integrates with Azure Synapse pipelines. Which approach should you use?

Easy
257

You have an Azure Data Factory pipeline that uses a Self-Hosted Integration Runtime (SHIR) to copy data from an on-premises Oracle database to Azure Blob Storage. The pipeline is failing with a connectivity error. You have verified that the SHIR is running and the network firewall allows outbound traffic to Azure. What is the most likely cause of the failure?

Medium
258

You have an Azure Synapse Analytics dedicated SQL pool that is used for reporting. You notice that the tempdb database is growing rapidly and causing queries to fail. Which two actions should you take to mitigate the issue? (Select two.)

Hard
259

You are building an Azure Stream Analytics job that reads JSON events from an Azure Event Hub and writes aggregated results to an Azure Synapse Analytics dedicated SQL pool. The events include a field named `EventTime` that is sometimes missing or malformed. You need the job to process only events with a valid `EventTime` and route invalid events to a separate output for later inspection. What should you do?

Medium
260

You are designing a data storage solution in Azure Synapse Analytics. You need to store large volumes of semi-structured data in a dedicated SQL pool. The data will be used for analytical queries that often filter on a date column and join on a customer ID column. You want to optimize query performance. Which two actions should you perform? (Choose two.)

Medium
261

You manage an Azure Synapse Analytics workspace with a dedicated SQL pool. The security team requires that all data stored in the dedicated SQL pool be encrypted with a customer-managed key (CMK) stored in Azure Key Vault. You need to configure transparent data encryption (TDE) to use the CMK. What should you do first?

Medium
262

A company uses Azure Synapse Analytics dedicated SQL pool. The data engineering team notices that queries against a large fact table are running slowly. The table uses round-robin distribution and has a columnstore index. The team wants to improve query performance without adding more resources. Which action should the team take?

Medium
263

You have a pipeline in Azure Data Factory that copies data from on-premises SQL Server to Azure Blob Storage. The pipeline fails with a 'Connection timed out' error. You have already verified that the Integration Runtime is running and the SQL Server firewall allows connections from the Integration Runtime. What should you check next?

Easy
264

You are a data engineer at a healthcare company. Your Azure Data Factory pipeline ingests sensitive patient records from an on-premises SQL Server into an Azure Data Lake Storage Gen2 account. The compliance team requires that all data be encrypted at rest with a customer-managed key (CMK) and that key rotation be audited. You need to configure the storage account to meet these requirements. What should you do?

Medium
265

You are designing a near-real-time data processing solution for a retail company. The source is a Kafka cluster on-premises. The target is an Azure Synapse Dedicated SQL Pool. The solution must handle up to 10,000 events per second with less than 5-minute latency. Which Azure service should you use to ingest the data?

Hard
266

You are designing a data processing solution in Azure Synapse Analytics. The solution must support both batch and streaming data ingestion. Which Azure service should you use to ingest streaming data into Synapse Analytics?

Easy
267

You are building an Azure Data Factory pipeline that processes files from Azure Blob Storage. The pipeline uses a Mapping Data Flow to transform the data and then writes the output to Azure Data Lake Storage Gen2. You need to ensure that the Data Flow can handle schema drift, where incoming files may have additional columns not present in the initial schema. What should you configure in the Data Flow?

Medium
268

You are designing a data processing solution for a global company. Data must be processed in near real-time and aggregated by region. You need to minimize latency for downstream consumers. Which Azure service should you use for stream processing?

Medium
269

You are developing an Azure Stream Analytics job that ingests telemetry from Azure Event Hubs and writes results to an Azure Synapse Analytics dedicated SQL pool. The job must compute a 5-minute tumbling window aggregation and write the aggregated rows to the dedicated SQL pool. You need to configure the output so that each window's aggregated rows are written efficiently. What should you do?

Medium
270

Which TWO actions should you take to secure sensitive data in Azure Data Lake Storage Gen2? (Choose two.)

Medium
271

Which Azure storage solution is best suited for storing large volumes of unstructured data, such as log files and media files, and supports both hierarchical namespace and POSIX-like access control lists?

Easy
272

You are optimizing an Azure Synapse Analytics dedicated SQL pool that experiences performance degradation during peak hours. You need to reduce query execution time by improving data distribution and reducing data movement. Which two actions should you take? (Choose two.)

Medium
273

Your company uses Azure Synapse Analytics to run a large-scale batch processing job every night. The job currently runs on a dedicated SQL pool and takes 4 hours. Management wants to reduce the runtime to under 2 hours without increasing cost. The job involves heavy compute operations with no data movement limitations. What should you do?

Hard
274

Which TWO Azure services can be used to implement a polyglot persistence architecture for an e-commerce application that requires both a relational database for orders and a document database for product catalogs?

Medium
275

A logistics company needs to store delivery tracking data that is updated frequently by multiple services. The solution must support transactions across multiple documents and provide real-time analytics. Which Azure service should you recommend?

Easy
276

You are monitoring an Azure Data Factory pipeline that copies data from Azure SQL Database to Azure Synapse Analytics. The pipeline occasionally fails with transient errors such as 'Cannot connect to SQL Database' during peak hours. You need to make the pipeline more resilient without manual intervention. What should you configure?

Hard
277

You are designing a data lake on Azure Data Lake Storage Gen2. The data includes customer PII that must be encrypted at rest using customer-managed keys. Which feature should you enable?

Medium
278

You need to design a data storage solution for a batch processing pipeline that processes petabytes of data daily. The data is stored in Parquet format and must be accessible by both Azure Databricks and Azure Synapse Analytics. Which storage solution should you recommend?

Easy
279

Refer to the exhibit. You have an Azure Data Factory pipeline that performs an incremental load from an Azure SQL Database source to a target Azure SQL Database. The pipeline uses a watermark column approach. After running the pipeline, you notice that the target table is empty. What is the most likely cause of this issue?

Hard
280

Which TWO factors should you consider when choosing between Azure SQL Database and Azure SQL Managed Instance for migrating a legacy application? (Choose two.)

Medium
281

Which TWO actions should you take to secure data in Azure Synapse Analytics dedicated SQL pool? (Choose two.)

Medium
282

Your Azure Synapse Analytics workspace uses serverless SQL pools for ad-hoc querying. Users report that queries are slow. You examine the execution plan and see that the query scans multiple partitions in the openrowset. What is the best way to improve performance?

Hard
283

You have an Azure Databricks notebook that processes a large Delta table and must be orchestrated from Azure Data Factory on a schedule. The notebook accepts two parameters, the source path and a run date. You need the pipeline to pass these values at runtime and to surface notebook failures as pipeline failures. Which activity configuration should you use?

Medium
284

You need to monitor resource utilization for an Azure Synapse Analytics dedicated SQL pool. Which Azure Monitor metric shows the percentage of allocated DWU being used?

Easy
285

You are designing a storage solution for a financial services company that uses Azure SQL Database. The database contains a table with sensitive customer data, including credit card numbers. Regulatory requirements mandate that the credit card numbers must be encrypted at rest and in use, and only authorized applications should be able to decrypt them. You need to implement a solution that allows encryption keys to be managed in Azure Key Vault. What should you use?

Medium
286

You are monitoring an Azure Synapse Analytics dedicated SQL pool and notice that some queries are experiencing high wait times due to concurrency slots being exhausted. You need to optimize the workload to reduce contention. Which three actions should you take? (Select three.)

Medium
287

You have an Azure Data Lake Storage Gen2 account that stores sensitive customer data. You need to prevent data exfiltration to unauthorized external IP addresses. Which TWO actions should you take?

Medium
288

Your team is troubleshooting slow query performance on a dedicated SQL pool in Azure Synapse Analytics. The query uses a hash-distributed fact table with 60 distributions. After reviewing the execution plan, you notice a high number of data moves. Which action would most likely reduce data movement?

Medium
289

Which THREE of the following are best practices for optimizing performance of Delta Lake tables in Azure Synapse Analytics? (Choose three.)

Hard
290

You need to assign permissions to a service principal so that it can write data to a specific container in Azure Data Lake Storage Gen2, but not delete blobs. The above JSON shows the built-in role 'Storage Blob Data Contributor'. The role includes delete permission in DataActions. What should you do?

Hard
291

You are configuring Azure Data Lake Storage Gen2 for a new data lake. You need to ensure that all data written to the 'raw' container is automatically encrypted at rest. Which feature should you enable?

Easy
292

You have an Azure Databricks notebook that processes data from a Delta table. The notebook runs slowly due to many small files. You need to optimize the Delta table for faster reads. Which Delta Lake operation should you run?

Easy
293

A company uses Azure Synapse Analytics dedicated SQL pool to store sales data. The sales table is partitioned by month and has a clustered columnstore index. Over time, the performance of queries filtering on a specific month has degraded. The data engineer suspects high rowgroup elimination. Which action should be taken to improve performance?

Hard
294

You have an Azure Data Factory pipeline that loads data from an on-premises SQL Server to Azure Synapse Analytics. The pipeline fails intermittently with network connectivity errors. You need to ensure reliable data transfer with minimal latency. Which solution should you recommend?

Hard
295

You are monitoring an Azure Data Factory pipeline that runs hourly. You notice that the pipeline has been failing intermittently with an error indicating 'Activity timeout'. Which Azure Monitor metric should you set an alert on to proactively detect such failures?

Easy
296

Which Azure service provides fully managed, serverless relational database capabilities for transactional workloads in a data storage solution?

Easy
297

You are building a data processing pipeline in Azure Synapse Analytics. The pipeline should read data from Azure Data Lake Storage Gen2 (Parquet files), apply transformations using a mapping data flow, and write the results to a dedicated SQL pool table. The source data contains personally identifiable information (PII). You need to mask the PII columns (e.g., email) using a data masking function within the data flow. Which transformation should you use?

Hard
298

You are building an Azure Synapse Analytics pipeline that uses a Mapping Data Flow to transform data from Azure Data Lake Storage Gen2. The data flow includes a derived column transformation that uses a custom expression to calculate a new field. You need to debug the data flow and preview the output at each transformation. Which feature should you use?

Medium
299

You are building a real-time dashboard that displays sales data from an Azure SQL Database. The dashboard must refresh every 30 seconds with minimal latency. You need to choose the appropriate Azure service for data processing and visualization. Which service should you use?

Medium
300

You are monitoring an Azure Synapse Analytics dedicated SQL pool. You notice that queries are occasionally queued due to concurrency limits. You need to reduce the impact of concurrency limits on query performance. What should you do?

Hard
301

Refer to the exhibit. A data engineer notices that the copy activity sometimes copies 0 rows despite reading 1 million rows. What is the most likely cause?

Hard
302

You need to store log files from multiple applications in a central location for long-term retention and occasional analysis. The data is rarely accessed after 30 days. Which storage solution should you use to minimize cost?

Easy
303

You are designing an Azure Synapse Analytics pipeline that uses a Mapping Data Flow to transform data from Azure Data Lake Storage Gen2. The data flow must handle schema drift, where new columns can appear in the source files over time. You need to ensure that the new columns are automatically included in the sink output without modifying the data flow. What should you do?

Medium
304

Which Azure service is primarily used for orchestrating data pipelines in a cloud-native ETL workflow?

Easy
305

You are monitoring an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline uses a self-hosted integration runtime. You notice that the copy activity sometimes takes much longer than expected, and you suspect network bottlenecks. You need to optimize the copy performance by adjusting the degree of parallelism. Which setting should you modify?

Medium
306

You are building an Azure Data Factory pipeline that must copy data from an on-premises SQL Server to Azure Blob Storage. The pipeline runs on a schedule every hour. You need to ensure that the copy activity can securely access the on-premises SQL Server. What should you configure?

Easy
307

You need to ensure that sensitive data stored in Azure SQL Database is encrypted at rest. Which feature should you enable?

Easy
308

You are using Azure Data Factory to copy data from an Azure SQL Database to an Azure Data Lake Storage Gen2 account. The copy activity is failing intermittently with a timeout error. You need to improve the throughput and reliability of the copy operation. What should you do?

Medium
309

You have an Azure Data Lake Storage Gen2 account that stores log files. You need to implement a data retention policy so that logs older than 90 days are automatically deleted. What should you use?

Easy
310

You are designing a data storage solution for a global e-commerce company. The company needs to store clickstream data from millions of users with high write throughput and low-latency reads for real-time analytics. The data is semi-structured and includes nested JSON objects. Which Azure data store should you recommend?

Medium
311

You are processing a large dataset in Azure Synapse Analytics using a dedicated SQL pool. You need to load data from Parquet files in Azure Data Lake Storage Gen2 into a staging table, then transform and load into a fact table. The fact table is partitioned by date. You want to maximize query performance and minimize data movement. Which technique should you use?

Hard
312

A company is designing a data lake solution on Azure Data Lake Storage Gen2. Data will be ingested from IoT devices at high frequency (every 5 seconds). Each device sends a JSON payload of 2 KB. The data must be stored in a hierarchical namespace and partitioned by date and device ID to optimize query performance. Which partition strategy should be used?

Medium
313

You are designing a streaming solution in Azure Synapse Analytics using the serverless SQL pool to query streaming data in real-time. The data is ingested via Azure Event Hubs and processed using Azure Stream Analytics. The output of Stream Analytics is written to Azure Data Lake Storage Gen2 in Delta Lake format. You need to ensure that the serverless SQL pool can query the latest data with minimal latency. Which approach should you use?

Medium
314

A logistics company needs to store shipment tracking events in Azure Cosmos DB. Events are written continuously throughout the day, and the most common query pattern retrieves all events for a specific shipment ID ordered by timestamp. The workload is write-heavy and must scale horizontally across partitions. Which partition key should you choose?

Easy
315

Your team is migrating an on-premises SQL Server data warehouse to Azure Synapse Analytics. The source data includes fact tables and dimension tables with complex relationships. You need to design the storage in Azure Synapse to minimize query latency for star schema queries. Which distribution and index strategy should you use for the fact table?

Hard
316

You need to transform semi-structured JSON data into a tabular format for analysis in Azure Synapse Analytics. The data is stored in ADLS Gen2. Which feature should you use to query the JSON data directly without loading it into a table?

Easy
317

A company runs a streaming pipeline using Azure Stream Analytics to ingest IoT data and output to Azure SQL Database. They notice that the output latency increases over time and eventually the job fails with a timeout error. What is the most likely cause?

Easy
318

You need to store historical sales data for 10 years with infrequent queries. The storage cost must be minimized while retaining the ability to query using Azure Synapse serverless SQL pool. Which storage tier should you use?

Easy
319

You are designing a security strategy for Azure Synapse Analytics. The solution must prevent users from accessing sensitive columns in a dedicated SQL pool, such as Social Security numbers, unless they have explicit permission. Which feature should you use?

Medium
320

Refer to the exhibit. You are reviewing the workload classifier configuration for an Azure Synapse Analytics dedicated SQL pool. You notice that the 'HeavyLoader' classifier has a queryExecutionTimeoutSeconds of 0. What is the implication of this setting?

Hard
321

A media company ingests high-definition video files into Azure Data Lake Storage Gen2. The files are uploaded once and then read by multiple analytics jobs for 48 hours, after which they are deleted. The company wants to optimize read performance and reduce latency for the analytics jobs. Which storage configuration should you recommend?

Medium
322

You are developing an Azure Data Factory pipeline that processes data from an on-premises SQL Server. The pipeline uses a self-hosted integration runtime. You need to ensure that the pipeline can handle schema changes in the source table without failing. What should you do?

Hard
323

Which THREE considerations are important when designing a table distribution strategy for an Azure Synapse Analytics dedicated SQL pool? (Choose three.)

Hard
324

You are designing a batch processing solution in Azure Synapse Analytics using pipelines. The solution must load data from multiple sources (Azure Blob Storage, Azure SQL Database, and REST API) into a dedicated SQL pool. After loading, you need to run a stored procedure to aggregate the data. Which two activities should you include in the pipeline? (Choose two.)

Medium
325

A data engineer needs to store log data from multiple applications in Azure. The data is append-only, heavily compressed, and queried infrequently. Cost minimization is critical. Which storage solution is best?

Easy
326

You are a data engineer for a retail company. The company uses Azure Data Lake Storage Gen2 to store raw transaction data partitioned by date. Each day, a folder is created with the format 'YYYY/MM/DD' containing thousands of small JSON files (each ~10 KB). An Azure Databricks job runs daily to read the previous day's folder, transform the data, and write to a Delta table for reporting. Over time, the job's execution time has increased from 15 minutes to over 2 hours. The job uses a cluster with 4 nodes (each 16 GB memory). Monitoring shows that the job spends most of its time in the 'listing files' stage. Which optimization should you implement to reduce the job duration?

Hard
327

You are designing a data storage solution for a retail company. The data includes transactional data that requires low-latency queries (under 10 milliseconds) and large historical data for analytics. The solution must minimize storage costs. Which approach should you recommend?

Easy
328

You are a data engineer for a retail company that stores sales data in an Azure Synapse Analytics dedicated SQL pool. You need to optimize query performance for a large fact table that is frequently joined with a much smaller dimension table. The queries often filter on a date column and aggregate sales amounts. Which technique should you implement to improve query performance?

Easy
329

Which TWO of the following Azure services can be used to orchestrate data pipelines that include data transformation?

Easy
330

Match each Azure data integration tool to its typical use case.

Medium
331

You need to ensure that data stored in Azure Data Lake Storage Gen2 is encrypted at rest using customer-managed keys. Which Azure service should you use to manage the keys?

Easy
332

Your organization has an Azure Synapse Analytics dedicated SQL pool that stores sensitive customer data. You need to ensure that only authorized users can access the data, and auditing must be enabled to track all access attempts. What should you do first?

Medium
333

You are monitoring Azure Stream Analytics job performance. The job is falling behind in processing real-time data. You notice that the SU (Streaming Unit) utilization is consistently at 90% or higher. What is the most appropriate action to improve throughput?

Easy
334

You need to monitor the performance of Azure Synapse Analytics dedicated SQL pool queries. Which Azure service should you use to identify long-running queries and resource bottlenecks?

Easy
335

You need to store semi-structured JSON data from a web application that requires low-latency reads and writes at a global scale. The data must be indexed automatically and support SQL-like queries. Which Azure data store should you use?

Easy
336

Refer to the exhibit. An Azure Policy is defined to enforce network security on storage accounts. What does this policy do?

Easy
337

You are designing data security for Azure Data Lake Storage Gen2. The requirement is to prevent data from being accessed by anyone outside the corporate network. Which feature should you enable?

Easy
338

You are developing an Azure Databricks notebook that processes streaming data from Azure Event Hubs. The notebook must write the processed data to a Delta table. You need to ensure that the stream can handle late data and update previously written records. Which Delta Lake feature should you use?

Medium
339

A company uses Azure Synapse Analytics dedicated SQL pool for data warehousing. They notice that queries against a large fact table are slow. The table is hash-distributed on ProductID, but many queries filter on OrderDate. What should the data engineer do to improve query performance?

Hard
340

You are designing a batch processing solution using Azure Databricks. The data source is a large Parquet dataset stored in Azure Data Lake Storage Gen2 (ADLS Gen2). The processing requires joining two datasets: one with 10 billion rows and another with 1 million rows. The cluster uses Photon runtime. Which optimization should you apply to minimize shuffle?

Hard
341

You have an Azure Databricks workspace that uses a managed resource group. The security team requires that all cluster nodes use no public IP addresses and that all outbound traffic goes through a firewall. What should you configure?

Medium
342

You are optimizing an Azure Synapse Analytics dedicated SQL pool. You need to reduce query execution time for large fact tables that are frequently joined with dimension tables. Which two actions should you perform? (Choose two.)

Hard
343

You are building an Azure Stream Analytics job that reads JSON events from an Azure Event Hub. Each event contains a nested array property named 'readings' with multiple sensor values. You need to output one row per sensor reading to an Azure Synapse Analytics dedicated SQL pool. The query must flatten the array. Which query syntax should you use?

Medium
344

Which TWO are benefits of using Azure Databricks Auto Loader for incremental data ingestion?

Easy
345

Which TWO Azure services can be used to perform data transformation in a data pipeline? (Select two.)

Easy
346

Which THREE components are required to implement a real-time data processing solution using Azure Stream Analytics?

Hard
347

Which of the following are valid methods to secure data at rest in Azure Data Lake Storage Gen2? (Choose two.)

Medium
348

You have an Azure Data Lake Storage Gen2 account that contains a container named raw. The container has a folder hierarchy with millions of small files. You need to optimize read performance for an Azure Databricks job that reads these files. You also need to minimize storage costs. What should you do?

Hard
349

You are designing a data processing solution for an e-commerce company. The company receives millions of clickstream events per hour from their website and needs to aggregate the data by product category and windowed time intervals for real-time dashboards. You need to minimize latency and cost. Which service should you use?

Medium
350

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline uses a self-hosted integration runtime. You notice that the copy activity fails intermittently with the error: 'Failure happened on 'Source' side. ErrorCode=SqlOperationFailed'. The on-premises SQL Server is under heavy load during business hours. What is the most likely cause?

Hard
351

You are a data engineer for a retail company that uses Azure Synapse Analytics. You have a dedicated SQL pool that contains a large fact table named SalesFact. Queries on SalesFact often filter by TransactionDate and join to a dimension table named Product. You notice that these queries perform poorly and sometimes spill to tempdb. You need to optimize the table design to improve query performance and reduce tempdb usage. What should you do?

Hard
352

You are designing a data processing pipeline in Azure Data Factory. The pipeline must copy data from Azure Blob Storage to Azure SQL Database and transform the data using a mapping data flow. The data flow includes a Derived Column transformation. What is the purpose of the Derived Column transformation?

Easy
353

You are implementing a data processing solution in Azure Databricks. The solution reads JSON files from Azure Data Lake Storage Gen2, performs complex transformations using PySpark, and writes the results to a Delta table. You need to ensure that the write operation is idempotent and can recover from failures without duplicating data. Which approach should you use?

Hard
354

You are implementing a Spark Structured Streaming job in Azure Databricks that reads from an Azure Event Hubs topic and writes to a Delta table. The job must handle late-arriving data up to 10 minutes and aggregate counts per device every 5 minutes. Which combination of settings should you use?

Hard
355

You need to orchestrate a data pipeline that includes a Python script and a Data Flow in Azure Synapse Analytics. The Python script must run before the Data Flow. Which activity should you use to run the Python script?

Easy
356

You are reviewing an Azure PowerShell script that sets permissions on a directory in Azure Data Lake Storage Gen2. The script sets a default ACL for a user on the path 'sales/2024/01/'. What is the effect of the -DefaultScope parameter?

Hard
357

A healthcare company stores patient records in Azure Data Lake Storage Gen2. The data must be encrypted at rest using customer-managed keys (CMK) stored in Azure Key Vault. The company also requires that the encryption keys are automatically rotated every 90 days. You need to configure the storage account to meet these requirements. What should you do?

Hard
358

You are troubleshooting an Azure Databricks job that writes data to Azure Data Lake Storage Gen2. The job fails with '403 Forbidden' error. The Databricks workspace uses a managed identity (system-assigned) for authentication. What should you verify?

Easy
359

You manage an Azure Data Lake Storage Gen2 account that stores sensitive financial data. The data must be encrypted at rest, and access must be audited. You need to ensure that encryption keys are managed by your organization and that all access attempts are logged. Which TWO actions should you take? (Choose two.)

Hard
360

You are using Azure Synapse Analytics dedicated SQL pool to process large fact tables. You need to improve query performance for joins between a large fact table and a small dimension table. The dimension table is less than 2 GB. What should you do?

Medium
361

A data engineer is designing a monitoring solution for Azure Data Factory pipelines. They need to be alerted when a pipeline run fails or when the duration exceeds a threshold. The solution must minimize cost and operational overhead. Which approach should they use?

Medium
362

Which TWO actions should you take to optimize query performance in Azure Synapse Analytics dedicated SQL pool when working with large fact tables?

Medium
363

You are optimizing an Azure Synapse Analytics dedicated SQL pool that contains a large fact table with over 1 billion rows. Queries frequently join this fact table with smaller dimension tables on a distribution key. You notice that many queries perform poorly due to data movement. You need to reduce data movement and improve query performance. Which two actions should you take? (Choose two.)

Hard
364

You manage an Azure Synapse Analytics dedicated SQL pool. A nightly ELT job loads a large fact table and then runs UPDATE statements on many rows. You observe that tempdb usage grows until the load fails. You need to reduce tempdb pressure during the update phase. What should you do?

Hard
365

An organization is using Azure Data Factory to ingest data from multiple on-premises SQL Server databases into Azure Synapse Analytics. They need to ensure that sensitive data is masked during ingestion before landing in the staging area. What is the best approach?

Hard
366

You are monitoring an Azure Data Lake Storage Gen2 account using Azure Monitor. You need to be alerted when the number of storage account requests exceeds 20,000 per hour. What is the most efficient way to set up this alert?

Medium
367

You are running a Spark job in Azure Synapse Analytics that reads from a Delta Lake table and performs multiple transformations. The job fails with an out-of-memory error on the executors. Which action should you take first to resolve the issue?

Easy
368

A retail analytics team stores Parquet files in Azure Data Lake Storage Gen2 partitioned by year, month, and day. Queries in Azure Synapse serverless SQL pools filter on a transaction date column, but performance is poor because the engine scans all files in the folder hierarchy. You need to reduce the amount of data scanned without changing the file layout. What should you implement?

Hard
369

A data engineering team uses Azure Stream Analytics to process real-time IoT data. They notice that the job's watermark delay is increasing over time, and the output is falling behind. The input is from Event Hubs with 10 partitions. The job uses a 5-minute hopping window with a 1-minute hop. What is the most likely cause?

Hard
370

Your team is developing a data processing solution that uses Azure Databricks to transform streaming data from Azure Event Hubs. The transformation includes joining the stream with a static reference table stored in Azure Data Lake Storage Gen2. You need to implement the join efficiently. Which approach should you use?

Easy
371

Refer to the exhibit. You have an Azure Synapse Analytics workspace. You need to ensure that data processing jobs can access the Data Lake Storage Gen2 account using a managed identity. What should you do?

Hard
372

You are a data engineer for a healthcare company. You have an Azure Data Lake Storage Gen2 account that stores sensitive patient data. You need to ensure that all access to the data is logged and that you can audit who accessed which files and when. You also need to minimize administrative effort. Which solution should you implement?

Medium
373

Your Azure Data Lake Storage Gen2 account stores sensitive data. You need to audit who accesses the data and when, and you want to send the audit logs to a Log Analytics workspace for analysis. What should you configure?

Easy
374

Match each Azure service to its primary purpose in a data engineering pipeline.

Medium
375

You are using Azure Databricks to process a large dataset stored in Azure Data Lake Storage Gen2. The data is in Parquet format and you need to optimize read performance for a query that filters on a specific column. What should you do?

Easy
376

You are developing an Azure Databricks notebook that processes JSON files stored in Azure Data Lake Storage Gen2. You need to read the files into a DataFrame and automatically infer the schema. Which code should you use?

Easy
377

You are implementing a data pipeline using Azure Data Factory. The source is an on-premises SQL Server database. Which Azure Data Factory component is required to connect to the on-premises data source?

Easy
378

You are monitoring an Azure Data Lake Storage Gen2 account that stores streaming data from IoT devices. You notice that query performance on the data in Parquet format is degrading over time. You need to improve query performance for both current and future data. Which TWO actions should you take?

Hard
379

You are designing a data storage solution for a retail company that needs to store transaction data that is frequently updated and requires strong consistency. The solution must support complex queries and joins across multiple tables. Which Azure data service should you recommend?

Medium
380

You are designing a streaming job in Azure Stream Analytics. The job needs to count the number of events per device type every 10 seconds. The input is from Event Hubs. Which query should you use?

Easy
381

You are using Azure Synapse Analytics dedicated SQL pool to run a query that joins a large fact table (10 billion rows) and a small dimension table (1 million rows). The query is slow. Which distribution strategy should you use for the dimension table to improve performance?

Medium
382

Your company uses Azure Data Factory to orchestrate data movement. You need to monitor pipeline runs across multiple factories and create a dashboard that shows success and failure rates over the past 30 days. What is the most efficient approach?

Medium
383

A manufacturing company uses Azure Data Lake Storage Gen2 to store IoT sensor data. The data arrives in JSON format with a nested structure. You need to transform the data into a tabular format for downstream analytics using Azure Synapse Pipelines. Which data flow transformation should you use?

Medium
384

You are implementing a data processing solution in Azure Synapse Analytics using Spark pools. The solution reads Parquet files from Azure Data Lake Storage Gen2, performs transformations, and writes the results to a dedicated SQL pool. You need to optimize the write performance to the dedicated SQL pool. Which technique should you use?

Medium
385

You are monitoring an Azure Data Factory pipeline that copies data from Azure Blob Storage to Azure SQL Database. The pipeline fails intermittently with the error: 'Operation on target SQL table failed: String or binary data would be truncated.' Which action should you take to resolve this issue?

Easy
386

You are a data engineer at a manufacturing company. You need to process sensor data from IoT devices that arrive in real time. The data is sent to Azure Event Hubs. You need to aggregate the data over 5-minute windows and store the results in Azure Data Lake Storage Gen2 in Parquet format. The solution should minimize cost and use serverless components. Which solution should you use?

Medium
387

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline is failing with a 'Gateway is offline' error. What is the most likely cause?

Easy
388

You are optimizing an Azure Synapse Analytics dedicated SQL pool. You need to reduce the amount of data read from storage during queries that filter on a date column. The fact table is partitioned by month on the date column. What should you do to improve query performance?

Easy
389

Which TWO options are recommended bulk loading methods for Azure Synapse SQL Pool? (Choose two.)

Hard
390

You are designing a data lake architecture for a healthcare company. The solution must support fine-grained access control at the file level, encryption at rest and in transit, and integration with Microsoft Purview for data lineage. Which storage solution should you recommend?

Hard
391

A company uses Azure Synapse Analytics dedicated SQL pool. They need to ensure that only users with a specific Azure AD group can query a particular schema. Which approach should they use?

Medium
392

You are a data engineer at a financial services company. You are developing a data processing pipeline that uses Azure Data Factory to copy transactional data from an Azure SQL Database to Azure Data Lake Storage Gen2. The pipeline runs daily and processes about 10 GB of data. You need to implement error handling for the pipeline. Specifically, if the copy activity fails due to a transient error, the pipeline should retry automatically. If the retry fails, the pipeline should log the error and send an email alert to the operations team. What should you do?

Easy
393

You are optimizing an Azure Synapse Analytics dedicated SQL pool that contains a fact table with 10 billion rows. Queries frequently join this fact table to a dimension table on a column that is not the distribution column of either table. You need to reduce data movement during these joins. Which two actions should you take? (Choose two.)

Hard
394

You are designing a data processing pipeline that ingests data from a REST API endpoint every hour. The API returns JSON data with a varying schema. You need to store the raw data in Azure Data Lake Storage Gen2 and later process it using Azure Databricks. Which file format should you use for the raw data storage?

Medium
395

You are writing a T-SQL query against a dedicated SQL pool in Azure Synapse Analytics. The query aggregates a fact table containing billions of rows by joining it to a small dimension table. You observe that the join produces a large amount of data movement and the query runs slowly. You need to reduce data movement for this recurring pattern. What should you do?

Medium
396

A company uses Azure Synapse Analytics with a dedicated SQL pool. They need to ensure that a team of data scientists can query all tables in the 'sales' schema but cannot modify any data or schema objects. Which role should the team be assigned?

Medium
397

You have an Azure Synapse Analytics dedicated SQL pool. You notice that some queries are taking longer than expected due to excessive data movement operations. You need to minimize data movement without changing the distribution columns. Which table design approach should you recommend?

Hard
398

You are reviewing a Spark job definition in Azure Synapse Analytics. The job aggregates sales data. The job runs successfully but takes longer than expected. You notice that dynamic allocation is disabled and the executor instances are fixed at 10. The cluster has a maximum of 20 nodes. What is the most likely reason for the slow performance?

Medium
399

You are using Azure Data Factory to copy data from an on-premises Oracle database to Azure Data Lake Storage Gen2. You need to ensure the copy activity can connect to the Oracle database without storing credentials in the pipeline JSON. What should you configure?

Medium
400

A data engineering team uses Azure Data Factory to load data from Azure SQL Database to Azure Data Lake Storage Gen2. They notice that the pipeline runs fail intermittently due to transient errors. They need to implement a retry policy with exponential backoff. What is the most efficient way to achieve this?

Hard
401

You are using Azure Purview to scan an Azure Data Lake Storage Gen2 account. After scanning, you notice that some files are not classified. What is the most likely reason?

Medium
402

Which TWO actions help optimize data storage costs in Azure Data Lake Storage Gen2?

Easy
403

You are designing a data lake architecture using Azure Data Lake Storage Gen2. The data will be ingested from multiple sources with varying schemas. You need to organize the data in a way that supports both batch and streaming analytics while maintaining data lineage. Which folder structure convention should you use?

Medium
404

You are designing a data processing solution using Azure Databricks with Delta Lake. The data is partitioned by date and ingested daily. You notice that the Delta table has many small files, causing slow read performance. Which strategy should you recommend to optimize the table for faster queries?

Hard
405

Match each Azure security feature to its description.

Medium
406

Which TWO options are correct approaches to handle schema drift in Azure Data Factory Mapping Data Flows?

Medium
407

A company is designing a data storage solution for a global application that requires low-latency reads and writes for user session data. The solution must support automatic failover across multiple Azure regions. Which TWO Azure services meet these requirements?

Medium
408

You have an Azure Data Factory pipeline that executes a stored procedure in Azure SQL Database. The pipeline fails with an error indicating that the stored procedure ran out of memory. What change should you make to the pipeline to resolve this?

Medium
409

You are developing a data processing solution in Azure Synapse Analytics. The solution must use a serverless SQL pool to query Parquet files stored in Azure Data Lake Storage Gen2. Which authentication method should you use to ensure that the queries use the identity of the caller and adhere to Azure role-based access control (RBAC) permissions?

Easy
410

Your organization uses Azure Synapse Analytics dedicated SQL pool to store sales data. You need to design a data loading process for a nightly batch that inserts new rows and updates existing rows based on the business key. The table has a clustered columnstore index. Which approach minimizes table fragmentation?

Medium
411

Which TWO options are valid methods to load data from on-premises SQL Server into Azure Synapse Analytics?

Easy
412

Which TWO actions should you take when monitoring Azure Data Lake Storage Gen2 to detect security threats?

Medium
413

You are processing streaming data from IoT devices using Azure Stream Analytics. The data includes temperature readings and device IDs. You need to calculate the average temperature per device over a 5-minute window, sliding every 1 minute. Which window function should you use?

Easy
414

You are developing a real-time data processing solution using Azure Stream Analytics. The input is an Azure Event Hubs stream with JSON data containing a 'timestamp' field. You need to output the average temperature per device every minute using a tumbling window. Which query should you use?

Easy
415

You need to configure encryption for an Azure SQL Database to protect data at rest. Which Azure service or feature should you enable?

Easy
416

You are designing a data processing solution in Azure Synapse Analytics. The solution must prevent unauthorized access to data at rest and in transit. Which combination of features should you implement?

Medium
417

Refer to the exhibit. You submit a Spark job in Azure Synapse Analytics using the Azure CLI. The job runs slowly during the shuffle phase. The input data is about 200 GB. Which configuration change would best improve performance for this shuffle-heavy workload?

Hard
418

A logistics company uses Azure Blob Storage to store shipping manifests as block blobs. The manifests are accessed frequently for the first 30 days, then rarely accessed for the next 60 days, and after 90 days they must be retained for seven years for compliance but are almost never accessed. You need to minimize storage costs while ensuring the data remains available for compliance audits. What should you do?

Easy
419

You deploy the Azure Security Center automation shown in the exhibit. What is the purpose of this automation?

Hard
420

Match each Azure Synapse Analytics component to its function.

Medium
421

You are using Azure Synapse Analytics serverless SQL pool to query Parquet files in Azure Data Lake Storage Gen2. The query returns fewer rows than expected. What should you check first?

Easy
422

Match the Azure service to its primary data processing use case. Drag each service on the left to the correct use case on the right. Services: Azure Databricks, Azure Stream Analytics, Azure Data Factory, Azure Synapse Analytics Use Cases: - Real-time event processing - Orchestration of ETL pipelines - Big data analytics with Spark - Enterprise data warehousing

Medium
423

You need to transform data in Azure Databricks using Apache Spark. The data is stored in Delta Lake format in Azure Data Lake Storage Gen2. Which method should you use to read the data into a Spark DataFrame?

Easy
424

You have an Azure Data Lake Storage Gen2 account that contains a container named raw with millions of small JSON files, each under 1 MB. A daily Azure Data Factory pipeline reads these files and writes them to a curated container as Parquet files. You notice that the pipeline runs slowly and you want to optimize read performance. What should you do first?

Hard
425

Which TWO actions can you take to optimize query performance in Azure Synapse Analytics dedicated SQL pool?

Medium
426

Drag and drop the steps to set up Azure Data Lake Storage Gen2 hierarchical namespace for a data lake into the correct order.

Medium
427

You are designing a data processing solution using Azure Databricks with Delta Lake. You need to ensure ACID transactions and schema enforcement. Which feature should you enable?

Hard
428

A company uses Azure Stream Analytics to process real-time data from IoT devices. They need to ensure that the output to Azure Synapse Analytics is optimized for high throughput and low latency. What should they configure in the Stream Analytics job?

Hard
429

Which TWO components are required to set up a streaming data pipeline using Azure Synapse Analytics? (Select two.)

Medium
430

You are monitoring an Azure Synapse Analytics dedicated SQL pool. You need to identify queries that are currently running and consuming the most resources. Which dynamic management view (DMV) should you query?

Medium
431

A company uses Azure Synapse Analytics dedicated SQL pool. They notice that queries against a large fact table are running slower over time. The table is hash-distributed on a date key and has a clustered columnstore index. Which action should you take to improve query performance?

Medium
432

A company uses Azure Synapse Analytics serverless SQL pool to query data in ADLS Gen2. Users report that queries against Parquet files are slow. What should you recommend to improve query performance?

Hard
433

You are designing a data storage solution for a global retail company that uses Azure Synapse Analytics dedicated SQL pool. The fact table is partitioned by date and contains 10 years of sales data. You need to implement a rolling window that keeps only the most recent 3 years of data while loading new daily data with minimal impact on concurrent queries. What should you do?

Hard
434

You have an Azure Data Lake Storage Gen2 account that stores parquet files. You need to ensure that files containing personally identifiable information (PII) are automatically classified and tagged. Which Azure service should you integrate?

Medium
435

You are designing a solution to store large amounts of log data that is written once and accessed rarely. The data must be retained for 7 years for compliance. After 30 days, the data should be moved to a lower-cost storage tier. After 1 year, the data should be archived. Which Azure Storage lifecycle management policy should you implement for an Azure Data Lake Storage Gen2 account?

Medium
436

You are designing a data lake architecture using Azure Data Lake Storage Gen2. You need to optimize query performance for Azure Synapse Analytics serverless SQL. Which three design considerations should you follow? (Choose three.)

Hard
437

Which TWO configurations are recommended to secure data processing in Azure Synapse Pipelines?

Easy
438

Your organization uses Azure SQL Database with Active Geo-Replication for disaster recovery. You need to ensure that all connections to the database use Microsoft Entra ID authentication and that access is audited. You also want to minimize the attack surface by disabling SQL authentication. What should you do?

Easy
439

A data engineer needs to store semi-structured JSON logs from multiple sources in Azure. The logs must be queryable using T-SQL and support schema-on-read. Which Azure service should be used?

Easy
440

You need to monitor an Azure Data Factory pipeline for failures and send an email notification when a pipeline run fails. Which Azure service should you use to create an alert based on the pipeline run metrics?

Easy
441

You are using Azure Synapse Analytics to process streaming data from Azure Event Hubs. The data must be written to a Delta Lake table in ADLS Gen2 with exactly-once semantics. Which processing engine should you use?

Easy
442

A healthcare organization needs to store electronic health records (EHR) in a format that supports schema flexibility and complex nested data. The solution must allow fast queries by patient ID and enable analytics with Azure Synapse. Which data store should you choose?

Easy
443

You need to monitor the performance of an Azure Synapse Analytics dedicated SQL pool. Which DMV should you query to find queries that are currently running and their execution status?

Easy
444

Which TWO features of Azure Databricks help manage data governance and security for sensitive data?

Easy
445

Your company uses Azure Blob Storage to store backups. You need to ensure that data is encrypted at rest using a customer-managed key stored in Azure Key Vault. Which feature should you enable?

Easy
446

You are building a streaming pipeline in Azure Stream Analytics that reads JSON events from an Azure Event Hub and writes to Azure Synapse Analytics. The events include a nested array of sensor readings. You need to flatten this array so each reading becomes a separate row. Which Stream Analytics feature should you use?

Medium
447

Refer to the exhibit. A user with Storage Blob Data Reader role on the container rawdata cannot list files under /2023/07/. What is the most likely reason?

Medium
448

You have a production pipeline in Azure Data Factory that copies data from an on-premises SQL Server to Azure Blob Storage using a self-hosted integration runtime. The pipeline fails intermittently with a 'Connection closed' error. The data volume is 50 GB per run. What should you first troubleshoot to resolve this issue?

Hard
449

You are optimizing an Azure Synapse Analytics dedicated SQL pool that processes large fact tables. You need to improve query performance for a common join between a fact table and a dimension table. The fact table is distributed using hash distribution on a column that is not the join key. The dimension table is small and replicated. You want to minimize data movement during the join. What should you do?

Hard
450

Your team is running a critical Azure Stream Analytics job that writes results to Azure SQL Database. Recently, the job has been failing with high latency and occasional data loss. You need to monitor the job's performance and set up alerts for when the watermark delay exceeds a threshold. What should you use?

Hard
451

A company is migrating its on-premises SQL Server data warehouse to Azure Synapse Analytics. They have a fact table with 2 billion rows and 30 columns. The table is frequently joined on CustomerID and filtered on OrderDate. What is the recommended table design?

Hard
452

Your Azure Data Lake Storage Gen2 account stores sensitive customer data. You need to ensure that data is encrypted at rest using customer-managed keys (CMK) and that access to the encryption key is logged. What should you do?

Hard
453

Which TWO of the following are supported storage options for use as a source in Azure Synapse Pipeline Copy Activity?

Medium
454

You are troubleshooting a failed Azure Synapse Pipeline execution. The pipeline uses a Copy activity to load data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The error indicates a 'Connection timeout' to the on-premises source. The Integration Runtime is Self-Hosted and has been running successfully for months. What is the most likely cause?

Medium
455

Refer to the exhibit. You are deploying an Azure Synapse Analytics workspace using an ARM template. The exhibit shows the encryption configuration. What is the effect of setting infrastructureEncryption to Enabled?

Medium
456

You are responsible for managing an Azure Data Lake Storage Gen2 account that stores parquet files for analytics. You need to implement a data retention policy that automatically deletes files older than 90 days in the 'logs' container. Additionally, you need to ensure that no data is lost due to accidental deletion; you want to be able to recover deleted files within 30 days. You also need to monitor the storage account for unusual access patterns. The solution must minimize administrative effort. What should you do?

Medium
457

You have a Data Factory pipeline that runs a U-SQL script in Azure Data Lake Analytics. The script processes terabytes of data and outputs to a CSV file. The pipeline is failing with the error: 'The job failed with UserError: Script execution failed.' You need to troubleshoot the issue. Which approach should you take first?

Hard
458

You are reviewing the ARM template snippet for an Azure Data Lake Storage Gen2 account. The template fails to deploy with an error that the encryption key cannot be accessed. What is the most likely cause?

Medium
459

You are designing a storage layer for an Azure Synapse Analytics dedicated SQL pool that ingests 4 TB of CSV files daily into a fact table. The files are landed in Azure Data Lake Storage Gen2 by an external ETL process. You need to load the data with the highest possible throughput while minimizing the load window. What should you do?

Medium
460

Refer to the exhibit. You are reviewing an ARM template for an Azure Data Lake Storage Gen2 account. Which of the following security best practices is violated in this template?

Hard
461

Refer to the exhibit. You are reviewing a Data Factory JSON definition. The factory has a user-assigned managed identity configured. However, the linked service to Azure Storage uses an account key. What security improvement should you recommend?

Medium
462

You need to implement column-level security in Azure Synapse Analytics to restrict access to salary information. Only users with the 'HRManager' role should see salary columns. Which feature should you use?

Easy
463

A data engineer needs to store semi-structured JSON logs from IoT devices. The data will be queried using SQL and must support high-throughput writes. Which Azure data store is most appropriate?

Easy
464

You manage an Azure Data Lake Storage Gen2 account used by an Azure Synapse Analytics workspace. You need to ensure that only authorized users can access data, and that all access attempts are logged for auditing. You configure Azure Active Directory (Azure AD) authentication and role-based access control (RBAC). Which additional feature should you enable to capture detailed access logs for compliance?

Hard
465

You are designing a data processing solution for a healthcare organization. The solution must process streaming data from IoT devices and store it in Azure Data Lake Storage Gen2. The data must be available for both real-time dashboards and historical analysis. You need to minimize operational overhead. What should you do?

Hard
466

Your organization uses Azure Synapse Analytics. You need to design a data transformation pipeline that processes streaming data from Azure Event Hubs, performs aggregations over a 5-minute tumbling window, and loads the results into a dedicated SQL pool table. Which Azure service should you use to implement the streaming transformation?

Medium
467

You are designing a data storage solution for IoT sensor data. The data is written thousands of times per second and requires low-latency reads for real-time dashboards. Which Azure storage solution should you use?

Easy
468

You are designing a data storage solution for a media company that stores video files in Azure Blob Storage. The company wants to optimize storage costs by automatically moving older files to cooler tiers. The files are accessed frequently for the first 30 days, then infrequently for the next 60 days, and rarely after that. You need to configure a lifecycle management policy. Which two actions should you include in the policy? (Choose two.)

Medium
469

You are building an Azure Stream Analytics job that ingests telemetry from Azure Event Hubs and writes aggregated results to an Azure Synapse Analytics dedicated SQL pool. The job must compute a five-minute tumbling window average per device and tolerate events that arrive up to three minutes late. During testing, you observe that events arriving after the window closes are silently dropped. You need to ensure late events are included in the correct window result. What should you configure in the Stream Analytics job?

Medium
470

You are designing a data lakehouse architecture in Azure using Delta Lake. The solution needs to process batch and streaming data from multiple sources, including IoT devices and CRM systems. You need to ensure data quality by enforcing schema validation and handling schema evolution. You also need to provide a unified catalog for querying. Which service should you use?

Medium
471

You are designing a storage solution for a financial services company. The solution must store large volumes of semi-structured JSON data in Azure Data Lake Storage Gen2. The data is accessed by Azure Databricks for batch processing and by Azure Synapse Analytics for interactive queries. The data must be organized for efficient partition elimination and must support atomic operations. You need to choose the appropriate file format and partitioning strategy. What should you do?

Hard
472

You manage an Azure Synapse Analytics dedicated SQL pool that stores a 4 TB fact table named FactSales. The table is currently distributed using ROUND_ROBIN and has a clustered columnstore index. Most analytical queries join FactSales to a much smaller dimension table DimProduct on ProductKey and then filter by DateKey. You need to redesign the physical storage to minimize data movement during these joins and improve query performance. What should you do?

Medium
473

You are designing a data processing solution in Azure Databricks to transform streaming data from Azure Event Hubs. The data must be aggregated in 1-minute tumbling windows and written to Azure Synapse Analytics. Which Spark API should you use?

Easy
474

A data engineer is setting up Azure Data Lake Storage Gen2 for a new project. The security requirement is to prevent direct access to the storage account from the internet while allowing access from a specific virtual network. Which network security feature should be enabled?

Easy
475

You are designing a disaster recovery plan for an Azure Synapse Analytics dedicated SQL pool. The primary region becomes unavailable. You need to fail over to a secondary region with minimal data loss. The recovery point objective (RPO) is 1 hour. What should you configure?

Hard
476

Which TWO are valid ways to process data in Azure Synapse Analytics?

Medium
477

You need to perform incremental data loading from Azure SQL Database to Azure Data Lake Storage Gen2 using Azure Data Factory. Which approach is the most efficient?

Easy
478

Your organization uses Azure Synapse Analytics serverless SQL pools to query data in Azure Data Lake Storage Gen2. You need to ensure that only users with specific Microsoft Entra ID roles can query the data. What should you configure?

Medium
479

You are designing a data processing solution in Azure Data Factory that must process files as they arrive in Azure Blob Storage. The solution must trigger a pipeline automatically when a new file is created, and then run a Databricks notebook to process the file. You need to configure the trigger and the pipeline activity. Which two actions should you perform? (Choose two.)

Medium
480

Your organization uses Azure Purview for data governance. You need to automatically scan an Azure Data Lake Storage Gen2 account and classify sensitive data such as credit card numbers and social security numbers. What should you configure?

Medium
481

You are designing a data processing solution in Azure Synapse Analytics. The solution must ensure that data at rest in a dedicated SQL pool is encrypted using customer-managed keys (CMK) stored in Azure Key Vault. The encryption should be enabled at the database level. What should you configure?

Medium
482

You have an Azure Synapse Analytics dedicated SQL pool. A nightly ELT process loads data into a staging table using PolyBase, then transforms and inserts it into a large fact table. You need to minimize data movement during the transformation step and ensure the fact table is optimized for large range scans. Which table distribution and index should you choose for the fact table?

Medium
483

You are implementing a streaming pipeline in Azure Stream Analytics that reads from an Azure Event Hub and writes aggregated results to an Azure Synapse Analytics dedicated SQL pool. The query groups events into 30-second windows. You need to ensure that the job can handle late-arriving events up to 2 minutes after the window closes without dropping them. What should you configure?

Medium
484

You are designing a data pipeline to ingest streaming data from IoT devices into Azure Synapse Analytics. The data must be available for querying with minimal latency, but you also need to handle spikes in throughput without data loss. Which service should you use as the ingestion layer?

Medium
485

You have an Azure Synapse Analytics serverless SQL pool that queries data in Azure Data Lake Storage Gen2. You need to ensure that only users with specific Microsoft Entra ID groups can access the data through the serverless SQL pool. What should you configure?

Medium
486

Drag and drop the steps to implement Slowly Changing Dimension (SCD) Type 2 in Azure Synapse Analytics dedicated SQL pool into the correct order.

Medium
487

You are monitoring an Azure Cosmos DB account using Azure Monitor. The 'Normalized RU Consumption' metric for a container is consistently above 90%. You need to ensure that the container can handle the load without throttling. What should you do?

Medium
488

You are tuning an Azure Stream Analytics job that reads from an Event Hub and writes to an Azure Synapse Analytics table. The job's SU% utilization is consistently at 90%. Which action would most likely reduce the SU% utilization?

Easy
489

A company is designing a data lake in Azure Data Lake Storage Gen2 (ADLS Gen2) to store IoT sensor data from millions of devices. The data is ingested in Parquet format, partitioned by date and device ID. The analytics team frequently queries the last 30 days of data for specific device types. Which partition strategy minimizes query cost and optimizes performance?

Medium
490

A company is planning to migrate an on-premises data warehouse to Azure Synapse Analytics dedicated SQL pool. The data warehouse contains a large fact table with billions of rows and several dimension tables. The company wants to optimize query performance and minimize data movement during joins. They need to choose an appropriate distribution type for the fact table. The fact table is frequently joined with dimension tables on a column that has high cardinality and is evenly distributed. What distribution type should they use?

Easy
491

You are monitoring an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline occasionally fails with a timeout error. You need to identify the cause of the failures and receive proactive alerts when similar issues occur. What should you do?

Easy
492

Your Azure Synapse Analytics dedicated SQL pool is experiencing performance degradation. Queries that previously completed in seconds now take minutes. You notice high queue wait times in sys.dm_pdw_exec_requests. What is the most likely cause?

Medium
493

Refer to the exhibit. You have an Azure Data Factory pipeline that copies trade data from Azure Blob Storage to Azure SQL Database. The pipeline runs every hour and truncates the destination table before each copy. However, users report that data is missing during the copy window. What is the most likely cause?

Medium
494

Your team is developing a data processing solution using Azure Databricks. The data is stored in Delta Lake format in Azure Data Lake Storage Gen2. You need to ensure that when multiple jobs concurrently write to the same Delta table, the operations are atomic and consistent. Which Delta Lake feature should you use?

Easy
495

Which THREE measures should you implement to monitor and optimize the performance of Azure Data Lake Storage Gen2?

Hard
496

You are designing a data processing solution in Azure Databricks. The data is stored in Azure Data Lake Storage Gen2 and you need to perform transformations using Apache Spark. The security requirements mandate that all data in transit must be encrypted and that the storage account must not be accessible from the public internet. What should you configure?

Easy
497

You are designing a data storage solution for real-time streaming data from IoT devices. The data must be stored in its original format for immediate processing and later transformed for analytics. Which Azure service should you use for raw data ingestion?

Easy
498

You are monitoring an Azure Data Factory pipeline that copies data from an Azure SQL Database to an Azure Data Lake Storage Gen2 account. The pipeline runs hourly. You notice that the copy activity sometimes takes much longer than expected. You need to identify the cause of the performance variability. Which action should you take first?

Medium
499

A company uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. They need to restrict a specific application's access to only write files in a particular directory without being able to read or list files. Which type of permission should be assigned?

Easy
500

You are designing a storage solution for a global application that requires low-latency reads and writes of JSON documents. The data model includes nested properties, and you need to query these properties efficiently. You also need to ensure the data is available in multiple regions with automatic failover. Which Azure service should you use?

Medium
501

Which THREE components are valid parts of the Microsoft Purview Data Map? (Choose THREE)

Hard
502

You have an Azure Synapse Analytics dedicated SQL pool that contains a large fact table. You need to minimize data movement during query execution for joins between the fact table and smaller dimension tables. What should you do?

Easy
503

Your company stores sensitive customer data in Azure SQL Database. You need to encrypt the data at rest and ensure that only your application can decrypt it, even from database administrators. What should you implement?

Medium
504

You are designing a data storage solution for a financial analytics platform. The platform ingests CSV files into Azure Data Lake Storage Gen2 and processes them with Azure Synapse Analytics serverless SQL pools. Queries frequently filter on a transaction date column and a region column, but the files are currently organized in a flat folder structure. You need to minimize the amount of data scanned by serverless SQL queries while keeping the files queryable using standard T-SQL OPENROWSET. What should you do?

Medium
505

Refer to the exhibit. A Bicep file is used to deploy an Azure Synapse Analytics workspace. What is the purpose of the 'purviewConfiguration' property?

Hard
506

Which THREE statements are true about partitioning in Azure Synapse Analytics dedicated SQL pool?

Hard
507

A company is designing a data storage solution for IoT device telemetry. Each device sends a JSON payload every second. The data must be stored in a way that supports real-time dashboards and long-term analytics with low latency. Which Azure data store should be used for the ingestion layer?

Easy
508

You are designing a security strategy for an Azure Data Lake Storage Gen2 account that stores sensitive financial data. The data must be encrypted at rest using customer-managed keys stored in Azure Key Vault. You also need to ensure that only specific Azure services can access the storage account. What should you do?

Medium
509

You need to transform data in Azure Synapse Analytics using a language that supports procedural logic and error handling. Which option should you use?

Easy

Frequently asked questions

What does the troubleshooting domain cover on the DP-203 exam?
troubleshooting questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 509 troubleshooting questions in the DP-203 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only troubleshooting questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.