Microsoft · Free Practice Questions · Last reviewed May 2026
18real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
35% of exam · 6 sample questions below
Your organization uses Azure Synapse Analytics dedicated SQL pool. You need to ensure that all data at rest in the SQL pool is encrypted using a customer-managed key stored in Azure Key Vault. What should you configure?
Implement Always Encrypted with column encryption keys stored in Azure Key Vault.
Configure Dynamic Data Masking to obfuscate sensitive data.
Enable Azure Storage Service Encryption with a customer-managed key.
Enable Transparent Data Encryption (TDE) with a customer-managed key in Azure Key Vault.
Transparent Data Encryption operates at the storage layer, encrypting data and log files at rest, and supports a customer-managed key held in Azure Key Vault — satisfying the stem's requirement that the dedicated SQL pool's data at rest be encrypted under organisational key control.
You have an Azure Databricks workspace that uses a managed resource group. The security team requires that all cluster nodes use no public IP addresses and that all outbound traffic goes through a firewall. What should you configure?
Configure service endpoints for Azure Storage and Azure Data Lake Storage.
Deploy the workspace in a VNet with forced tunneling enabled and a firewall.
Deploying into a customer-managed VNet with forced tunnelling routes all cluster node egress through your firewall appliance, satisfying the outbound inspection requirement. Secure cluster connectivity (no public IP) is enabled alongside, so nodes hold no public addresses. This meets both constraints the security team specified.
Apply network security groups (NSGs) to the subnet that restrict outbound traffic.
Enable Azure Private Link for the Databricks workspace.
Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to grant a service principal read and write access to a specific directory without granting access to the parent directories. What should you use?
Assign the Storage Blob Data Contributor role at the directory level using RBAC.
Use a managed identity and assign it to the directory.
Create a stored access policy on the directory.
Set ACLs on the directory with default ACLs for the service principal.
Default ACLs apply only to new child items created within the directory, not to the directory itself or to existing objects, so they cannot grant the service principal read and write access to the target directory directly. This option is tempting because default ACLs are designed to propagate permissions to future files and subdirectories, making them correct when the requirement is to control access for newly created content rather than the existing directory.
Refer to the exhibit. You are reviewing the workload classifier configuration for an Azure Synapse Analytics dedicated SQL pool. You notice that the 'HeavyLoader' classifier has a queryExecutionTimeoutSeconds of 0. What is the implication of this setting?
Queries classified as 'HeavyLoader' will wait indefinitely for resources.
The configuration is invalid; queryExecutionTimeoutSeconds must be greater than 0.
Queries classified as 'HeavyLoader' will not have a timeout.
A queryExecutionTimeoutSeconds value of zero disables the timeout entirely, so 'HeavyLoader' queries run until completion or manual cancellation. This satisfies the stem's scenario by confirming that the classifier imposes no execution time limit, unlike classifiers with a positive value that terminate queries after the specified seconds.
Queries classified as 'HeavyLoader' will timeout immediately.
You have an Azure Synapse Analytics serverless SQL pool. You need to monitor the number of queries that are currently executing. Which dynamic management view should you query?
sys.dm_resource_governor_workload_groups
sys.dm_exec_query_stats
sys.dm_exec_requests
sys.dm_exec_requests shows currently executing requests in the serverless SQL pool, including state, command, and session ID.
sys.dm_exec_sessions
Your team is using Azure Synapse Analytics to process sensitive customer data. You need to ensure that column-level security is applied to a specific table so that only users with a certain role can view certain columns. Which feature should you use?
Column-level security (CLS)
Column-level security uses GRANT and DENY statements on individual columns, restricting which roles can read specified fields within the table. This directly enforces the requirement that only users holding a certain role view particular columns.
Row-level security (RLS)
Azure Purview data policies
Dynamic data masking (DDM)
Want more Secure, monitor, and optimize data storage and data processing practice?
Practice this domain46% of exam · 6 sample questions below
You are designing a batch processing solution using Azure Databricks. The data source is a large Parquet dataset stored in Azure Data Lake Storage Gen2 (ADLS Gen2). The processing requires joining two datasets: one with 10 billion rows and another with 1 million rows. The cluster uses Photon runtime. Which optimization should you apply to minimize shuffle?
Broadcast the smaller table (1 million rows) to all worker nodes.
Broadcasting the smaller table avoids shuffling the large table, significantly reducing data movement.
Increase the cluster size to reduce shuffle overhead.
Create bucketed tables on the join key for both datasets.
Use Delta Lake and optimize file layout with OPTIMIZE command.
You are running a Spark job in Azure Synapse Analytics that reads from a Delta Lake table and performs multiple transformations. The job fails with an out-of-memory error on the executors. Which action should you take first to resolve the issue?
Enable checkpointing to truncate the lineage.
Decrease the number of partitions to reduce overhead.
Increase the executor memory setting in the Spark configuration.
Executor out-of-memory errors arise when each executor's JVM heap cannot hold the partition data during transformations. Raising spark.executor.memory gives those executors more heap, directly relieving the constraint. Partition tuning or skew handling may follow, but increasing memory is the quickest first action.
Use the cache() action on intermediate DataFrames.
You are optimizing a Spark DataFrame transformation in Azure Synapse Analytics. The DataFrame has 20 columns and 100 million rows. You notice that the job is slow due to many small files being written to the output. Which two actions can you take to reduce the number of output files? (Choose two.)
Use coalesce() to reduce the number of partitions without a shuffle.
coalesce() merges partitions into the requested count without a full shuffle, so each task writes fewer, larger files. On a 100-million-row DataFrame this directly reduces the small-file problem while avoiding the network cost of repartitioning.
Enable caching on the DataFrame before writing.
Apply bucketing on a column to group data.
Increase the number of partitions using repartition() with a larger number.
Use repartition() with a smaller number of partitions.
repartition() with a smaller partition count performs a full shuffle that redistributes rows evenly, so each output task writes one larger file. This reduces the number of small files, though it costs more than coalesce() because of the shuffle.
You are building a data pipeline that uses Azure Data Factory to copy data from a REST API to Azure Blob Storage. The REST API returns JSON data in pages of 1000 records each. The total number of records is 50,000. Which activity or feature should you use to loop through the pages?
Use a ForEach activity to iterate over a fixed number of pages.
Use a Lookup activity to retrieve the total number of pages and then use a ForEach.
Use an Until activity to loop until the API returns no more pages.
Use a Copy activity with pagination rules enabled in the source.
Copy activity pagination rules let the source iterate REST API pages automatically using properties such as NextPageUrl or absolute URLs, retrieving all 50,000 records without a ForEach loop. This satisfies the stem's requirement to traverse 1,000-record pages efficiently.
You are designing a streaming job in Azure Stream Analytics. The job needs to count the number of events per device type every 10 seconds. The input is from Event Hubs. Which query should you use?
SELECT DeviceType, COUNT(*) FROM Input GROUP BY DeviceType, SessionWindow(second, 10, 30)
SELECT DeviceType, COUNT(*) FROM Input GROUP BY DeviceType, TumblingWindow(second, 10)
TumblingWindow(second, 10) partitions events into fixed, non-overlapping 10-second intervals, satisfying the stem's requirement to count per device type every 10 seconds. Grouping by DeviceType alongside the window produces one count per device type per interval, which sliding or session windows cannot guarantee.
SELECT DeviceType, COUNT(*) FROM Input GROUP BY DeviceType, HoppingWindow(second, 10, 1)
SELECT DeviceType, COUNT(*) FROM Input GROUP BY DeviceType, SlidingWindow(second, 10)
You are using Azure Synapse Analytics to process streaming data from Azure Event Hubs. The data must be written to a Delta Lake table in ADLS Gen2 with exactly-once semantics. Which processing engine should you use?
Azure Databricks with Structured Streaming
Azure Databricks Structured Streaming natively supports Delta Lake sinks with idempotent writes and checkpointing, delivering exactly-once semantics when consuming Event Hubs. This satisfies the stem's exactly-once requirement, which plain Spark or Synapse streaming cannot guarantee without additional transactional handling.
Azure Synapse serverless SQL pool
Azure Synapse Pipeline with Mapping Data Flow
Azure Stream Analytics
Want more Develop data processing practice?
Practice this domain19% of exam · 6 sample questions below
A company is designing a data lake solution on Azure Data Lake Storage Gen2. Data will be ingested from IoT devices at high frequency (every 5 seconds). Each device sends a JSON payload of 2 KB. The data must be stored in a hierarchical namespace and partitioned by date and device ID to optimize query performance. Which partition strategy should be used?
Use Azure SQL Database with clustered columnstore index on date and device ID.
Organize folders as /YYYY/MM/DD/DeviceID/ in ADLS Gen2 and use file naming that includes timestamp.
ADLS Gen2 hierarchical namespace supports true directory semantics, so /YYYY/MM/DD/DeviceID/ paths let partition pruning skip irrelevant folders during queries. Date-first ordering suits time-range filters, while DeviceID narrows per-device scans, and timestamped filenames preserve ingestion order within each partition.
Use Azure Table Storage with PartitionKey set to date and RowKey set to device ID.
Use Azure Cosmos DB with partition key on (date, device ID) and TTL for data retention.
You are designing a near-real-time analytics pipeline for a retail company. Transaction data is generated in Azure SQL Database and must be replicated to Azure Synapse Analytics (dedicated SQL pool) with less than 5 minutes latency. The source table has 50 million rows and 200 columns, but only 30 columns are needed for analytics. Which approach should you recommend?
Use Azure SQL Database Change Tracking and push changes to Azure Event Hubs, then use Azure Stream Analytics to write to Synapse.
Enable Change Data Capture (CDC) on the source table and use Azure Data Factory with a 1-minute tumbling window to copy changes into Synapse.
CDC captures only changed rows, and ADF can run frequently to meet latency target.
Use Azure Synapse PolyBase to directly query the source SQL database every 5 minutes.
Schedule a full copy of the entire table every 5 minutes using Azure Data Factory.
A data engineer needs to store semi-structured JSON log files from a web application. Each log entry is about 1 KB. The logs are rarely queried (once a month) and must be retained for 7 years for compliance. The solution must minimize storage cost. Which storage option should be used?
Store the logs in Azure SQL Database as a table.
Store the logs in Azure Files share.
Store the logs in Azure Blob Storage with cool access tier.
Blob Storage cool tier is low-cost for infrequent access, suitable for logs.
Store the logs in Azure Cosmos DB with a JSON container.
Which TWO of the following are supported storage options for use as a source in Azure Synapse Pipeline Copy Activity?
Azure Data Lake Storage Gen2
ADLS Gen2 is a supported source.
Azure Analysis Services
Azure Cognitive Search
Azure Purview
Azure Blob Storage
Azure Blob Storage is a supported source.
You are an administrator for an Azure Synapse Analytics dedicated SQL pool. You execute the T-SQL statements shown in the exhibit. The external table 'dbo.Orders' is created. Which statement about querying this external table is true?
Querying the external table automatically imports data into a round-robin distribution.
The table cannot be queried until the data is imported into the dedicated SQL pool.
You can query the external table using standard T-SQL SELECT statements.
Dedicated SQL pools expose external tables through the PolyBase engine, so standard T-SQL SELECT statements work against them exactly as with internal tables. The external table definition supplies the schema and the LOCATION pointing to the data in Azure Data Lake Storage, satisfying the requirement to query without loading.
You must first create a PolyBase external table before querying.
A company is designing a data storage solution for IoT device telemetry. Each device sends a JSON payload every second. The data must be stored in a way that supports real-time dashboards and long-term analytics with low latency. Which Azure data store should be used for the ingestion layer?
Azure SQL Database
Azure Blob Storage
Azure Event Hubs
Event Hubs ingests millions of telemetry events per second with low latency, buffering the stream for downstream consumers. This satisfies the real-time dashboard requirement while retaining data for long-term analytics, unlike batch-oriented stores such as Blob Storage or Azure SQL Database.
Azure Data Lake Storage
Want more Design and implement data storage practice?
Practice this domainThe DP-203 exam has 50 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 3 domains: Secure, monitor, and optimize data storage and data processing, Develop data processing, Design and implement data storage. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Microsoft DP-203 exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.