Courseiva

CCNA Secure, monitor, and optimize data storage and data processing Questions

75 of 159 questions · Page 1/3 · Secure, monitor, and optimize data storage and data processing · Answers revealed

1
Multi-Selectmedium

You are securing an Azure Data Lake Storage Gen2 account that contains sensitive data. Which TWO of the following should you implement to protect data from unauthorized access?

Select 2 answers
A.Configure ACLs to grant least privilege to users and groups
B.Use private endpoints to restrict access to the storage account
C.Set the default ACL to allow read access for all authenticated users
D.Enable CORS rules to allow only specific origins
E.Enable large file shares on the storage account
AnswersA, B

POSIX-style ACLs on Data Lake Storage Gen2 apply at directory and file level, enforcing least-privilege access for individual users and groups independently of broad role assignments. This granular permission model directly prevents unauthorised reads by limiting each principal to only the paths they require.

Why this answer

Option A is correct because Azure Data Lake Storage Gen2 uses POSIX-style access control lists (ACLs) at both the directory and file level, and configuring ACLs to grant least privilege ensures users and groups only receive the specific read/write/execute permissions they require, directly preventing unauthorized access to sensitive data. Option B is correct because private endpoints assign a private IP address from your virtual network to the storage account, removing exposure to the public internet and restricting access to only clients within the approved virtual network or connected networks. Option C is incorrect because setting a default ACL that allows read access to all authenticated users violates least privilege and would broaden, not restrict, access to sensitive data.

Option D is incorrect because CORS rules only control which web origins can make cross-origin browser requests to the service; they do not authenticate users or prevent unauthorized direct access. Option E is incorrect because enabling large file shares only increases the maximum capacity and file size limits of the file share, and has no effect on access control or authorization.

Exam trap

DP-203 often tests the confusion between network-layer controls (private endpoints, firewall rules) and identity-layer controls (RBAC, ACLs) — candidates who pick CORS or large file shares mistake non-security features for access controls.

2
MCQmedium

Your company uses Azure Data Lake Storage Gen2 and needs to implement a data retention policy that automatically deletes files older than 90 days in a specific container. What should you use?

A.Azure Data Factory pipeline with a Delete activity scheduled to run daily.
B.Azure Policy with a deny effect for files older than 90 days.
C.Azure Storage lifecycle management rule with a filter for the container and a delete action after 90 days.
D.Azure Purview data lifecycle policy.
AnswerC

Lifecycle management rules evaluate blob age and apply actions automatically, so a rule scoped to the container with a delete action after 90 days enforces the retention policy without custom code. This matches Azure Data Lake Storage Gen2's native tiering and expiry capability.

Why this answer

Azure Storage lifecycle management policies are the native, server-side mechanism for automatically transitioning or deleting blobs based on age. A rule can be scoped with a prefix filter matching the specific container and configured with a delete action after 90 days, executing automatically without any external orchestrator. This is the correct, cost-effective, and operationally simple solution for ADLS Gen2 retention.

Exam trap

DP-203 often tests whether candidates reach for custom orchestration (Data Factory) or governance tooling (Azure Policy, Purview) when the native, purpose-built feature — Storage lifecycle management — is the correct and simplest answer.

How to eliminate wrong answers

Option A is wrong because an Azure Data Factory pipeline with a Delete activity is a custom, scheduled workaround that adds compute cost, requires orchestration monitoring, and is unnecessary when native lifecycle rules exist. Option B is wrong because Azure Policy enforces resource-level governance (e.g., allowed SKUs, tags) and cannot evaluate or delete individual blob files based on age — it has no per-object lifecycle capability. Option D is wrong because Azure Purview is a data governance/catalog service; it does not execute storage deletion actions on a schedule.

3
MCQmedium

You are a data engineer at a financial services company. You have an Azure Data Lake Storage Gen2 account named finlake that stores sensitive transaction data in Parquet files. You need to ensure that data is encrypted at rest using a customer-managed key stored in Azure Key Vault, and that the key is automatically rotated every 90 days. You also need to be able to revoke access to the data immediately if the key is compromised. What should you do?

A.Enable infrastructure encryption (double encryption) on the storage account and store the keys in a hardware security module (HSM) without configuring rotation.
B.Configure Azure Storage encryption with customer-managed keys in Azure Key Vault, set the key rotation policy to 90 days, and grant the storage account access to the key via a managed identity.
C.Enable Azure Storage Service Encryption with Microsoft-managed keys and configure a lifecycle management policy to rotate keys every 90 days.
D.Use Azure Disk Encryption with BitLocker keys stored in Azure Key Vault and apply the encryption to the storage account's underlying disks.
AnswerB

This option uses customer-managed keys stored in Azure Key Vault, which allows you to control rotation and revoke access by disabling or deleting the key. Configuring a rotation policy in Key Vault automates rotation. Granting the storage account a managed identity with access to the key vault enables secure key usage without storing secrets.

Why this answer

To encrypt data at rest in Azure Data Lake Storage Gen2 with customer-managed keys, you must configure the storage account to use a key stored in Azure Key Vault. A managed identity assigned to the storage account grants access to the key vault. You can set a rotation policy in Key Vault to automatically rotate the key every 90 days.

If the key is compromised, disabling or deleting it immediately revokes access to the data.

Exam trap

The trap here is assuming that Microsoft-managed keys can be rotated on a custom schedule or that revoking them is possible; only customer-managed keys provide that level of control.

4
MCQeasy

You need to secure data at rest for an Azure Data Lake Storage Gen2 account that contains sensitive financial data. Which configuration should you enable to ensure that data is encrypted using a customer-managed key stored in Azure Key Vault, and that access to the key is logged?

A.Enable Azure Storage encryption with Microsoft-managed keys
B.Implement client-side encryption using Azure Key Vault
C.Enable infrastructure encryption for double encryption
D.Configure Azure Storage encryption with customer-managed keys in Azure Key Vault and enable Key Vault logging
AnswerD

Customer-managed keys in Azure Key Vault replace Microsoft-managed keys for storage encryption, and enabling Key Vault logging records every key access. Together these satisfy both the CMK encryption requirement and the key-access auditing constraint for the financial data.

Why this answer

Azure Storage encryption with customer-managed keys in Key Vault provides control and logging. Option A is wrong because Microsoft-managed keys are the default but do not provide customer control. Option B is wrong because client-side encryption requires managing keys on the client side.

Option C is wrong because infrastructure encryption adds a second layer but does not use customer-managed keys.

5
MCQmedium

Your organization uses Azure Data Lake Storage Gen2 for a data lake. You need to prevent accidental deletion of data by enabling a soft delete policy. Which configuration is required?

A.Apply an Azure Resource Manager lock.
B.Configure Azure Backup for the storage account.
C.Enable blob versioning.
D.Enable blob soft delete on the storage account.
AnswerD

Blob soft delete retains deleted blobs for a configurable retention period, letting you recover data removed accidentally. Enabling it at the storage account level protects all containers, satisfying the requirement to prevent accidental deletion in the Data Lake Storage Gen2 account.

Why this answer

Azure Data Lake Storage Gen2 supports blob soft delete, which protects against accidental deletion by retaining deleted blobs for a specified retention period. Option A is incorrect because Azure Resource Manager locks prevent the deletion or modification of the storage account itself, not the data within it. Option B is incorrect because Azure Backup is designed for backing up VMs, SQL databases, and other workloads, not for managing soft delete of blobs.

Option C is incorrect because blob versioning preserves previous versions of blobs, but it does not prevent deletion of the current version; soft delete is specifically for recovery from accidental deletion.

6
MCQeasy

You store sensitive data in Azure Data Lake Storage Gen2. You need to ensure that only members of a specific security group can read the data, while other users in the organization must not have access, even if they have the Storage Blob Data Reader role at the storage account level. What should you use?

A.Azure Private Link with a private endpoint for the storage account.
B.Access control lists (ACLs) on the directory or file, with entries for the security group and a mask.
C.Shared access signatures (SAS) scoped to the container.
D.Azure role-based access control (RBAC) at the storage account level.
AnswerB

ACLs in Azure Data Lake Storage Gen2 allow you to grant permissions to specific Microsoft Entra ID security groups at the directory or file level. By setting an ACL entry for the group and a mask, you can ensure that only members of that group have read access, even if others have broader RBAC roles. ACLs are evaluated together with RBAC to determine effective permissions.

Why this answer

Access control lists in Azure Data Lake Storage Gen2 provide file and directory-level permissions that can be assigned to Microsoft Entra ID security groups. By granting read access to the specific group and using a mask, you ensure that only group members can read the data. RBAC roles at the account level are too broad and cannot restrict access to a subset of users.

Exam trap

The trap here is assuming that account-level RBAC roles or network controls like Private Link can restrict access to a specific group, when only POSIX-style ACLs on directories or files provide that granularity.

7
MCQmedium

You manage an Azure Data Lake Storage Gen2 account containing a large volume of JSON files. Users report that direct read operations from the data lake are slow, and you observe high egress costs. You need to optimize read performance and reduce cost for analytical queries that frequently filter on a specific timestamp column and select a subset of columns. What should you do?

A.Convert the files to Parquet format and partition the data by the timestamp column.
B.Enable Azure Storage analytics logging and review the logs to identify slow queries.
C.Increase the number of partitions in the Azure Synapse Analytics dedicated SQL pool that reads the data.
D.Move the data to a premium block blob storage account with a higher throughput tier.
AnswerA

Parquet is a columnar format that enables column pruning and predicate pushdown, reducing I/O and cost. Partitioning by the timestamp column further limits the data scanned when queries filter on that column. This directly addresses slow reads and high egress by minimizing the amount of data transferred and processed.

Why this answer

Converting JSON to Parquet reduces storage size and enables columnar reads, while partitioning by the frequently filtered timestamp column allows the query engine to skip irrelevant data. Together, these changes minimize the data scanned and transferred, improving read performance and lowering egress costs. Other options either do not address the root cause or introduce unnecessary expense.

Exam trap

The trap here is assuming that simply moving data to a higher-performance storage tier or increasing compute resources will solve read performance and cost issues, without addressing the inefficient data format and lack of partitioning.

8
MCQhard

Your team uses Azure Databricks for data processing. You need to implement a cost-control strategy that automatically terminates idle clusters after 30 minutes of inactivity, but allows users to override this policy for specific workloads that require long-running clusters. What is the most efficient approach?

A.Instruct all users to set auto-termination to 30 minutes on each cluster they create.
B.Configure a global auto-termination setting in the Azure Databricks workspace that terminates all clusters after 30 minutes of inactivity.
C.Use Azure Policy to enforce a tag that triggers a function to terminate idle clusters.
D.Create a cluster policy that enforces auto-termination with a default of 30 minutes, but allows users to override the value for specific clusters.
AnswerD

A cluster policy sets auto-termination at 30 minutes by default while permitting overrides, so idle clusters terminate automatically yet long-running workloads can extend the timeout. This satisfies both the cost-control and override requirements without manual intervention.

Why this answer

Azure Databricks cluster policies let administrators define constraints and defaults for cluster configuration, including auto-termination. A cluster policy can set a default auto-termination of 30 minutes while allowing users to override the value (within policy-permitted bounds) for specific long-running workloads. This centralizes governance without blocking legitimate exceptions.

Exam trap

DP-203 often tests the difference between a hard global limit (which removes flexibility) and a policy-based default with override (which balances governance and flexibility) — candidates who pick the global setting miss the override requirement.

How to eliminate wrong answers

Option A is wrong because relying on users to manually set auto-termination is not enforced and will inevitably be missed, defeating the cost-control objective. Option B is wrong because a global workspace setting that terminates all clusters after 30 minutes removes the ability for users to override for long-running workloads, violating the requirement. Option C is wrong because Azure Policy operates on Azure resource metadata and cannot directly terminate idle Databricks clusters; it is the wrong control plane and would require custom automation that is far less efficient than a native cluster policy.

9
MCQmedium

You have an Azure Databricks workspace that processes sensitive data. The security team requires that all access to the workspace be authenticated using Microsoft Entra ID and that all API calls be audited. Which configuration should you implement?

A.Configure workspace to use Microsoft Entra ID authentication and enable diagnostic settings for audit logs.
B.Enable VNet injection and configure network security groups.
C.Deploy Azure Private Link and disable public access.
D.Configure personal access tokens for API access and enable cluster logs.
AnswerA

Configuring the workspace for Microsoft Entra ID authentication enforces identity-based access, while diagnostic settings stream audit logs to a Log Analytics workspace or storage account. Together these satisfy both constraints: Entra ID authentication for all access and auditable API calls.

Why this answer

To meet both requirements—Microsoft Entra ID authentication and audited API calls—you must configure the workspace to use Entra ID authentication and enable diagnostic settings to send audit logs to a destination like Log Analytics. This combination ensures identity-based access and a record of API activity.

Exam trap

The trap is focusing on network security (Private Link, VNet injection) when the question explicitly asks for authentication and auditing; candidates may overlook the need for diagnostic settings to capture API calls.

How to eliminate wrong answers

Option B is wrong because VNet injection and NSGs address network isolation, not authentication or API auditing. Option C is wrong because Private Link and disabling public access improve network security but do not enforce Entra ID authentication or provide API audit logs. Option D is wrong because personal access tokens are a separate authentication mechanism that does not use Entra ID, and cluster logs do not capture all API calls for auditing.

10
MCQeasy

Your organization uses Azure Data Factory to orchestrate data pipelines. You need to ensure that sensitive data is not exposed in pipeline logs. What should you configure?

A.Store connection strings in Azure Key Vault.
B.Enable 'Secure output' on pipeline activities.
C.Set a retention policy for pipeline logs.
D.Use data flow debug logs with session logs.
AnswerB

Enabling 'Secure output' prevents activity output from being written to Data Factory monitoring logs, directly satisfying the requirement that sensitive data must not appear in pipeline logs. Unlike 'Secure input', which masks only input values, this setting suppresses the logged output itself, ensuring credentials or personal data captured during execution remain hidden.

Why this answer

Enabling 'Secure output' on pipeline activities prevents sensitive data from being written to Azure Data Factory logs. Option A is incorrect because while storing connection strings in Azure Key Vault is a security best practice, it does not prevent sensitive data from appearing in pipeline logs. Option C is incorrect because setting a retention policy for logs controls how long logs are kept, but does not prevent sensitive data exposure.

Option D is incorrect because data flow debug logs are for debugging and do not mask sensitive data in pipeline logs.

11
MCQmedium

Your company uses Azure Purview for data governance. You need to ensure that sensitive data in Azure Data Lake Storage Gen2 is automatically detected and classified. What should you configure in Purview?

A.Apply sensitivity labels to the storage account using Microsoft Purview Information Protection.
B.Enable Microsoft Defender for Cloud's data sensitivity discovery.
C.Use Azure Policy to enforce tagging of resources containing sensitive data.
D.Create a scan rule set that includes built-in classification rules for sensitive data types.
AnswerD

A scan rule set containing built-in classification rules tells Purview which sensitive data types to detect during scans of Data Lake Storage Gen2. Classification then applies automatically, satisfying the requirement for automatic detection and classification of sensitive data.

Why this answer

In Microsoft Purview, you can create scan rule sets that include built-in classification rules to automatically detect sensitive data types during scanning. Option A is incorrect because sensitivity labels are applied after classification, not for detection. Option B is incorrect because Microsoft Defender for Cloud's data sensitivity discovery is a different feature; Purview itself handles classification.

Option C is incorrect because Azure Policy enforces compliance rules, not data classification at the file level.

12
MCQhard

You are monitoring an Azure Synapse Analytics dedicated SQL pool and notice that queries against a large fact table are slow. The table is distributed using hash distribution on a column that has a high number of nulls. You need to improve query performance. What should you do?

A.Create a clustered columnstore index on the table and rebuild it.
B.Partition the table on the column with many nulls to isolate the null values.
C.Change the distribution to round-robin to evenly distribute the data.
D.Change the distribution column to a column with high cardinality and even distribution, such as a surrogate key.
AnswerD

Hash distribution on a column with many nulls causes skew because all nulls are placed in the same distribution. Choosing a column with high cardinality and even distribution (e.g., a surrogate key) ensures data is evenly spread across distributions, improving query parallelism and reducing data movement for joins on that column.

Why this answer

Hash distribution on a column with many nulls leads to data skew because all nulls hash to the same distribution. This causes one distribution to have more data and processing load, slowing down queries. Choosing a distribution column with high cardinality and even distribution, such as a surrogate key, ensures balanced data across distributions and improves query performance.

Exam trap

The trap here is focusing on indexing or partitioning when the root cause is distribution skew due to a poor choice of hash distribution column.

13
MCQeasy

You need to grant a data analyst read access to a specific folder in an Azure Data Lake Storage Gen2 account. The analyst must not be able to read other folders in the same container. You want to follow the principle of least privilege. What should you use?

A.The Storage Blob Data Contributor role assigned at the storage account scope.
B.A POSIX access control list (ACL) on the target folder, with the analyst's Azure AD identity granted read and execute.
C.A storage account shared access signature scoped to the container.
D.A stored access policy on the container that allows read operations.
AnswerB

Data Lake Storage Gen2 supports POSIX-style ACLs at the directory and file level. Granting the analyst's Azure AD identity read and execute on the specific folder, without granting permissions on the parent or sibling folders, enforces least privilege. Execute permission on the parent directories is needed for traversal, but it does not grant read access to their contents.

Why this answer

POSIX ACLs in Data Lake Storage Gen2 allow permission entries at the directory and file level, so you can grant an identity read and execute on one folder while leaving sibling folders inaccessible. This is the native mechanism for fine-grained, least-privilege access. Broad role assignments and container-scoped SAS or policies grant access to the whole container and are therefore unsuitable here.

Exam trap

The trap here is confusing role-based access control, which is typically scoped at the account or container level, with directory-level POSIX ACLs that can isolate a single folder.

14
MCQhard

You are designing a data processing solution using Azure Databricks. The solution must use Delta Lake for ACID transactions and must optimize storage costs by automatically compacting small files. Which feature should you enable?

A.Run OPTIMIZE command in a scheduled job.
B.Set retention duration for vacuum to 7 days.
C.Enable Z-order indexing on the Delta table.
D.Enable auto-optimize on the Delta table.
AnswerD

Auto-optimize automatically compacts small files during writes, removing the need for manual OPTIMIZE jobs while preserving Delta Lake ACID guarantees. This directly satisfies the storage-cost constraint by reducing many small files into larger ones without extra orchestration.

Why this answer

Auto-optimize (also called Optimized Writes / auto compaction) is a Delta Lake table property that automatically compacts small files during writes, eliminating the need for manual OPTIMIZE jobs. It combines with optimized writes to reduce the small-file problem at ingestion time, directly lowering storage and query costs. This is the native, built-in feature designed for the stated requirement.

Exam trap

DP-203 often tests the distinction between automatic table maintenance features (auto-optimize, optimized writes) and manual commands (OPTIMIZE, VACUUM, ZORDER), so candidates who see 'OPTIMIZE' and assume it means automatic compaction pick the wrong answer.

How to eliminate wrong answers

Option A is wrong because running OPTIMIZE in a scheduled job is a manual workaround that requires orchestration and only compacts periodically, not automatically — it does not satisfy 'automatically compacting.' Option B is wrong because VACUUM retention controls how long old file versions are kept before deletion; it has nothing to do with compaction of small files. Option C is wrong because Z-order indexing is a data-clustering technique applied via OPTIMIZE ZORDER BY to improve data-skipping on frequently filtered columns — it does not automatically compact small files.

15
MCQeasy

You need to monitor the performance of your Azure Synapse Analytics dedicated SQL pool. Which metric should you use to identify queued queries due to concurrency limits?

A.Queued queries
B.DWU percentage
C.Active queries
D.Memory percentage
AnswerA

The queued queries metric counts requests waiting for a concurrency slot in the dedicated SQL pool. Rising values indicate that active queries have exhausted available slots, directly identifying concurrency-limit queuing rather than resource pressure such as DWU saturation or tempdb usage.

Why this answer

The 'Queued queries' metric in Azure Synapse Analytics dedicated SQL pool specifically counts queries waiting for resources due to concurrency limits, making it the direct indicator of queuing caused by insufficient concurrency slots. Monitoring this metric helps identify when the workload exceeds available concurrency and queries are being held in the queue.

Exam trap

DP-203 often tests the distinction between metrics that indicate resource saturation (DWU percentage, memory percentage) and those that specifically indicate concurrency queuing, so candidates must know that 'Queued queries' is the direct signal.

How to eliminate wrong answers

Option B is wrong because DWU percentage measures the utilization of Data Warehouse Units (compute resources) and does not directly indicate queued queries. Option C is wrong because Active queries shows queries currently executing, not those waiting in the queue. Option D is wrong because Memory percentage reflects memory consumption and is not the primary metric for concurrency-related queuing.

16
Multi-Selectmedium

Which TWO Azure services can be used to monitor and analyze query performance in Azure Synapse Analytics dedicated SQL pool?

Select 2 answers
A.SQL Data Sync
B.Azure Policy
C.Azure Advisor
D.Azure Monitor with Log Analytics
E.Dynamic Management Views (DMVs)
AnswersD, E

Dedicated SQL pool emits query execution metrics to Azure Monitor, and Log Analytics stores and queries them via KQL. This satisfies the monitoring requirement by correlating DMV-level query performance data across the pool, which neither storage nor networking services provide.

Why this answer

Azure Monitor with Log Analytics (D) is correct because dedicated SQL pool emits diagnostic logs and metrics (e.g., DmvLogs, QueryStoreRuntimeStatistics, ExecRequests) that can be routed to a Log Analytics workspace, where KQL queries and workbooks are used to monitor and analyze query performance over time. Dynamic Management Views (E) are correct because T-SQL DMVs such as sys.dm_pdw_exec_requests, sys.dm_pdw_exec_sessions, sys.dm_pdw_request_steps, and sys.dm_pdw_sql_requests expose real-time execution details, wait statistics, and step-level timings for queries running in the dedicated SQL pool. SQL Data Sync (A) is incorrect because it is a data synchronization service for Azure SQL Database and SQL Managed Instance, not a query performance monitoring tool for Synapse dedicated SQL pools.

Azure Policy (B) is incorrect because it enforces governance and compliance rules on resources, not query performance analysis. Azure Advisor (C) is incorrect because it provides general best-practice recommendations (cost, reliability, performance) at the resource level, not detailed query-level monitoring or analysis for dedicated SQL pools.

Exam trap

DP-203 often tests the distinction between real-time diagnostic tools (DMVs) and historical monitoring services (Azure Monitor), causing candidates to overlook DMVs as a monitoring option or confuse Azure Advisor's recommendations with actual performance monitoring.

17
MCQmedium

You are designing a data lake architecture using Azure Data Lake Storage Gen2. You need to implement a least-privilege security model. Which authorization mechanism should you use for granular control?

A.Use storage account keys for access.
B.Use Azure RBAC roles at the storage account level.
C.Use POSIX-like access control lists (ACLs).
D.Use shared access signatures (SAS) with stored access policies.
AnswerC

POSIX-like ACLs grant per-file and per-directory permissions to individual security principals, exceeding what coarse role assignments allow. This satisfies the least-privilege requirement by scoping read, write and execute rights at the folder or file level within Azure Data Lake Storage Gen2.

Why this answer

POSIX-like access control lists (ACLs) in Azure Data Lake Storage Gen2 provide file- and directory-level granular permissions, enabling least-privilege access down to individual users or groups. They are the recommended mechanism when you need fine-grained control beyond what Azure RBAC roles at the container or storage account level can provide. ACLs support both access ACLs (for current access) and default ACLs (inherited by new child items).

Exam trap

DP-203 often tests the confusion between RBAC (coarse, management-plane, container-level) and ACLs (fine-grained, data-plane, file/directory-level) — candidates pick RBAC because it is the more familiar Azure authorization model, but it cannot deliver least-privilege at the file level.

How to eliminate wrong answers

Option A is wrong because storage account keys grant full administrative access to the entire storage account — the opposite of least privilege — and cannot be scoped to individual files or directories. Option B is wrong because Azure RBAC roles at the storage account level operate at the management plane and are too coarse for file-level granularity; they grant broad permissions like 'Storage Blob Data Reader' across the whole account or container. Option D is wrong because shared access signatures (SAS) with stored access policies are time-bound, delegated access tokens for external sharing or temporary access, not a mechanism for persistent, granular, identity-based least-privilege control within the data lake.

18
Multi-Selecteasy

Which THREE best practices should be followed when designing a data lake in Azure Data Lake Storage Gen2 for optimal performance?

Select 3 answers
A.Disable hierarchical namespace to improve performance.
B.Use a deep directory structure with many subfolders.
C.Use Parquet file format for analytics workloads.
D.Use a naming convention that avoids special characters and high cardinality.
E.Partition data by date to enable partition elimination.
AnswersC, D, E

Parquet's columnar, compressed layout lets analytical engines read only the referenced columns and skip irrelevant row groups, sharply cutting I/O against the stem's optimal-performance requirement. Row-based formats force full-row deserialisation regardless of projection, so Parquet directly satisfies the constraint that analytics workloads over Data Lake Storage Gen2 must minimise scanned data.

Why this answer

Option C is correct because columnar formats like Parquet compress data efficiently and let analytics engines such as Spark, Synapse, and Databricks read only the needed columns, dramatically reducing I/O and improving query performance. Option D is correct because ADLS Gen2 (and the underlying Blob REST/ABFS APIs) handles simple, predictable names more efficiently; special characters can break path parsing and high-cardinality names (e.g., GUIDs or timestamps in every filename) defeat caching, listing, and partition pruning. Option E is correct because partitioning data by date (e.g., year=/month=/day=) aligns with how engines like Spark, Hive, and Synapse perform partition elimination, so queries scanning a single day skip all other partitions instead of reading the whole dataset.

Option A is wrong because enabling the hierarchical namespace is what makes ADLS Gen2 a true data lake with directory-level operations, atomic renames, and better performance for analytics; disabling it reverts to flat Blob storage semantics. Option B is wrong because deep, heavily nested directory trees increase metadata and listing overhead and slow down operations; a shallow, well-partitioned structure is preferred.

Exam trap

The trap is that candidates assume disabling the hierarchical namespace improves performance (it actually removes key optimizations) and that deeper folder hierarchies are better, when flat, date-partitioned layouts are the recommended design.

19
MCQeasy

You are configuring security for an Azure Data Lake Storage Gen2 account. You need to ensure that users can only access files and folders for which they have explicit permissions, and that permissions are enforced at the file and folder level. What should you enable?

A.Access control lists (ACLs) only
B.Shared access signatures (SAS) only
C.Azure RBAC and access control lists (ACLs)
D.Azure role-based access control (Azure RBAC) only
AnswerC

To enforce file and folder level permissions in Azure Data Lake Storage Gen2, you must use both Azure RBAC and ACLs. RBAC grants the user or service principal access to the storage account or container, while ACLs provide fine-grained permissions on individual files and folders. This combination ensures that users can only access the data for which they have explicit permissions, meeting the requirement.

Why this answer

Azure Data Lake Storage Gen2 uses a combination of Azure RBAC and POSIX-like ACLs to secure data. RBAC controls access at the management and container level, while ACLs provide fine-grained permissions at the file and folder level. To ensure users can only access files and folders for which they have explicit permissions, both must be configured.

RBAC alone is too coarse, ACLs alone lack the authentication context, and SAS tokens are not identity-based for per-user permissions.

Exam trap

The trap here is thinking that ACLs alone or RBAC alone can provide complete file-level security, when in fact they are complementary.

20
MCQeasy

You are running an Azure Stream Analytics job that reads from an Event Hub and writes to a Power BI dataset. The job is falling behind and processing latency is increasing. What should you do to improve performance?

A.Increase the number of Streaming Units (SUs) allocated to the job.
B.Use a reference data input to filter events.
C.Change the output to Azure Blob Storage instead of Power BI.
D.Decrease the size of events sent to the Event Hub.
AnswerA

Streaming Units govern the compute and parallelism available to the job; when throughput exceeds current capacity, partitions queue and latency climbs. Scaling SUs raises processing bandwidth, directly addressing the backlog constraint without altering the query or Event Hub configuration.

Why this answer

Azure Stream Analytics scales throughput by allocating Streaming Units (SUs), which represent a bundle of CPU, memory, and network resources. When a job falls behind, increasing SUs provides more parallel processing capacity, directly reducing latency. This is the standard vertical scaling mechanism for Stream Analytics and is the correct first response to performance degradation.

Exam trap

DP-203 often tests the misconception that changing the output or input configuration will improve performance, when the primary scaling lever for Stream Analytics is increasing Streaming Units (SUs).

How to eliminate wrong answers

Option B is wrong because reference data is used for enriching or filtering events based on static or slowly changing data, not for improving throughput; it adds a join operation that can actually increase latency. Option C is wrong because changing the output to Blob Storage does not address the root cause of the job falling behind; it merely changes the sink and may even reduce performance if the job is output-bound, but the question asks about improving performance, not changing output. Option D is wrong because decreasing event size might reduce network and processing overhead per event, but it does not increase the job's processing capacity and is not a direct scaling action; moreover, it may not be feasible if the event schema is fixed.

21
MCQeasy

You need to ensure that data in an Azure Data Lake Storage Gen2 account is encrypted at rest using a customer-managed key. Which feature should you configure?

A.Azure Key Vault integration with Storage Service Encryption
B.Azure Information Protection
C.Azure Storage Service Encryption with Microsoft-managed keys
D.Azure Disk Encryption
AnswerA

Configuring Azure Key Vault integration lets Storage Service Encryption use a customer-managed key stored in Key Vault rather than a Microsoft-managed key, giving the organisation control over the key lifecycle. This satisfies the requirement for encryption at rest with a customer-managed key.

Why this answer

Customer-managed keys (CMK) for Azure Storage are implemented by integrating the storage account with Azure Key Vault (or Managed HSM) and configuring the account's encryption key to reference a key stored there. Storage Service Encryption (SSE) then uses that key to wrap the account's data encryption key, giving the customer control over key lifecycle and revocation.

Exam trap

The trap is confusing data-classification services (Azure Information Protection) or disk-level encryption (ADE) with storage-account encryption; only Key Vault integration with SSE provides customer-managed keys for ADLS Gen2.

How to eliminate wrong answers

Option B is wrong because Azure Information Protection is a data-classification and labeling service for documents and emails, not a storage-at-rest encryption mechanism for Data Lake Storage Gen2. Option C is wrong because Microsoft-managed keys are the default and do not satisfy the requirement for customer-managed keys — the customer has no control over rotation or revocation. Option D is wrong because Azure Disk Encryption applies BitLocker/DM-Crypt to IaaS VM OS and data disks, not to PaaS storage accounts such as ADLS Gen2.

22
MCQeasy

Your organization uses Azure Data Lake Storage Gen2 and needs to prevent accidental deletion of data by enabling soft delete. You also need to ensure that deleted blobs are recoverable for 30 days. What should you configure?

A.Enable blob snapshots and set them to expire after 30 days.
B.Use Azure Backup to create daily backups of the storage account.
C.Enable container soft delete with a retention period of 30 days.
D.Enable blob soft delete and set retention period to 30 days.
AnswerD

Blob soft delete retains deleted blobs for a configurable retention period, satisfying the 30-day recoverability constraint. Setting the retention period to 30 days ensures accidental deletions remain restorable within that window, directly meeting the stated requirement.

Why this answer

Blob soft delete enables recovery of deleted blobs within a specified retention period (30 days in this case). Option A is incorrect because blob snapshots are point-in-time copies that require manual management and do not provide automatic recovery of deleted blobs. Option B is incorrect because Azure Backup is designed for virtual machines and other Azure resources, not for blob-level recovery in Data Lake Storage Gen2.

Option C is incorrect because container soft delete deletes entire containers, not individual blobs.

Exam trap

A common trap is confusing blob soft delete with container soft delete. Container soft delete protects entire containers, while blob soft delete protects individual blobs. For this question, blob soft delete is required.

23
Multi-Selecthard

Which THREE security features are available in Azure Data Lake Storage Gen2 to protect data at rest and in transit? (Choose three.)

Select 3 answers
A.Azure Storage firewalls and virtual network rules
B.Azure Information Protection
C.Encryption at rest using Storage Service Encryption (SSE)
D.Azure ADLS Gen2 supports HTTPS for data in transit.
E.Azure Policy
AnswersA, C, D

Storage firewalls and virtual network rules restrict network access to the storage account, blocking public internet traffic and permitting only approved subnets or IP ranges. This protects data in transit by enforcing network-level perimeter controls on Azure Data Lake Storage Gen2.

Why this answer

Option A (Azure Storage firewalls and virtual network rules) is correct because ADLS Gen2, built on Azure Storage, lets you restrict network access to the storage account by configuring firewall rules and allowing specific virtual networks and subnets, thereby protecting data at rest from unauthorized network access. Option C (Encryption at rest using Storage Service Encryption, SSE) is correct because Azure Storage automatically encrypts data at rest using 256-bit AES encryption, and this applies to ADLS Gen2 accounts, protecting data at rest. Option D (HTTPS for data in transit) is correct because ADLS Gen2 supports secure transfer over HTTPS/TLS, ensuring data is encrypted in transit between clients and the storage service.

Option B (Azure Information Protection) is not a native ADLS Gen2 data-at-rest or in-transit protection feature; it is a separate classification and labeling service for documents and emails. Option E (Azure Policy) is a governance and compliance tool for enforcing organizational rules, not a direct data-at-rest or in-transit encryption/network protection feature of ADLS Gen2.

Exam trap

DP-203 often tests the confusion between governance services (Azure Policy, Azure Information Protection) and native storage security controls, tricking candidates into selecting services that 'sound' security-related but do not directly protect ADLS Gen2 data.

24
Multi-Selecthard

Which THREE metrics should you monitor for an Azure Synapse Analytics dedicated SQL pool to ensure optimal performance?

Select 3 answers
A.tempdb usage
B.DWU usage
C.Queued queries
D.Login failures
E.Total storage size
AnswersA, B, C

tempdb usage in a dedicated SQL pool indicates spill-to-disk from insufficient memory grants during large sorts, joins and aggregations. Monitoring it reveals queries exceeding granted memory, which degrade performance and signal the need for statistics updates or query tuning.

Why this answer

Monitoring tempdb usage (A) is essential because dedicated SQL pool queries spill to tempdb for sorts, hash joins, and other operations, and high tempdb utilization can indicate insufficient memory or poorly designed queries that degrade performance. DWU usage (B) is a key metric because it reflects how much of the provisioned Data Warehouse Unit capacity is being consumed; consistently high DWU usage signals that the pool may need scaling or workload tuning. Queued queries (C) should be monitored because a growing queue indicates that concurrency limits or resource contention are delaying query execution, directly affecting performance.

The unmarked options do not belong: login failures (D) is a security/connectivity metric rather than a performance indicator, and total storage size (E) relates to capacity planning and cost, not to query performance optimization.

Exam trap

The trap is confusing capacity metrics (like storage size) with performance metrics. Candidates might think login failures are a performance issue, but they are security-related. Also, DWU usage is sometimes overlooked as it directly measures compute utilization.

25
MCQhard

Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to implement a security strategy that allows users to read only specific folders within a container. Which authorization method should you use?

A.Storage account shared key
B.Azure RBAC roles (e.g., Storage Blob Data Contributor) at the container level
C.Shared access signatures (SAS) with folder-level permissions
D.Access control lists (ACLs) on the folder
AnswerD

ACLs provide POSIX-style permissions at file and directory level, so individual folders within a container can grant read access independently. This satisfies the requirement for folder-scoped authorisation, which container-level RBAC roles cannot achieve because they apply to the whole container.

Why this answer

ACLs (Access Control Lists) in Azure Data Lake Storage Gen2 can be applied to individual folders, enabling granular read permissions. Option A is incorrect because a storage account shared key grants full access to the entire account. Option B is incorrect because Azure RBAC roles like Storage Blob Data Contributor apply at the container level, affecting all folders within.

Option C is incorrect because shared access signatures (SAS) can be scoped to a container or a file, but not to a specific folder within a container.

26
MCQmedium

You have an Azure Synapse Analytics dedicated SQL pool that contains a large fact table named FactSales. The table is partitioned by date and has a clustered columnstore index. You notice that queries filtering on a specific date range are slow. You need to improve query performance for these queries. What should you do?

A.Rebuild the clustered columnstore index on the FactSales table.
B.Create a nonclustered index on the date column.
C.Ensure that the queries use a predicate on the partitioning column so that partition elimination occurs.
D.Switch the table distribution to round-robin.
AnswerC

Partition elimination allows the query optimizer to skip partitions that do not contain relevant data. If queries filter on the partitioning column (e.g., date), and the predicate is sargable, the engine can scan only the necessary partitions. This directly reduces I/O and improves performance for date-range queries.

Why this answer

Partition elimination is a key performance feature in dedicated SQL pools. When queries filter on the partitioning column with a sargable predicate, the engine can prune partitions and read only the relevant data. This reduces the amount of data scanned and speeds up queries.

Ensuring that queries are written to take advantage of partition elimination is the most direct solution for slow date-range queries.

Exam trap

The trap here is assuming that index maintenance or distribution changes will fix slow date-range queries, when the actual issue is often lack of partition elimination.

27
MCQeasy

You need to monitor the performance of an Azure Stream Analytics job in real time. Which Azure service should you use to track the job's resource utilization (e.g., SU % utilization) and set up alerts when the job is approaching its capacity?

A.Azure Monitor
B.Azure Advisor
C.Microsoft Sentinel
D.Azure Log Analytics
AnswerA

Azure Monitor collects Stream Analytics job metrics including SU % utilisation and supports metric alerts on thresholds. It satisfies the stem's requirement to track resource utilisation in real time and notify when the job approaches capacity, unlike log-only or billing-focused services.

Why this answer

Azure Monitor is the correct service for real-time monitoring of Azure Stream Analytics jobs, providing metrics such as SU % utilization, watermark delay, and input/output events. It allows you to create alert rules based on these metrics to notify when the job approaches capacity limits.

Exam trap

DP-203 often tests the distinction between Azure Monitor (metrics and alerts) and Log Analytics (log queries), causing candidates to choose Log Analytics for real-time metric monitoring.

How to eliminate wrong answers

Option B is wrong because Azure Advisor provides best-practice recommendations, not real-time performance monitoring or alerting. Option C is wrong because Microsoft Sentinel is a SIEM/SOAR solution for security analytics, not job performance monitoring. Option D is wrong because Azure Log Analytics is a log query and analysis service; while it can store logs, it does not natively provide the real-time metric tracking and alerting for SU utilization that Azure Monitor does.

28
Multi-Selectmedium

Which TWO Azure services can be used to monitor Azure Data Factory pipeline runs and set up alerts?

Select 2 answers
A.Log Analytics
B.Microsoft Sentinel
C.Azure Policy
D.Azure Monitor
E.Azure Advisor
AnswersA, D

Log Analytics stores Azure Data Factory diagnostic logs, and alert rules defined against that workspace trigger notifications on pipeline run failures or duration thresholds. This satisfies the monitoring and alerting requirement for Data Factory pipeline runs.

Why this answer

Azure Monitor (option D) is the native platform for collecting metrics and activity/diagnostic logs from Azure Data Factory and for creating alert rules (metric alerts, log search alerts, activity log alerts) that trigger on pipeline run failures or durations, so it directly satisfies the monitoring-and-alerting requirement. Log Analytics (option A) is the service where ADF diagnostic logs are stored and queried with KQL, and it powers log search alerts (via Azure Monitor) on pipeline run records such as PipelineRun and ActivityRun, making it a valid way to monitor runs and set alerts. Microsoft Sentinel (B) is a SIEM/SOAR for security analytics, not pipeline operational monitoring; Azure Policy (C) enforces governance/compliance rules rather than monitoring runs; and Azure Advisor (E) only provides best-practice recommendations, not run monitoring or alerting.

Exam trap

DP-203 often tests the distinction between monitoring/alerting services (Azure Monitor, Log Analytics) and governance/security services (Azure Policy, Microsoft Sentinel, Azure Advisor), so candidates may incorrectly select Azure Policy for alerting or Microsoft Sentinel for pipeline monitoring.

29
MCQeasy

You have an Azure Synapse Analytics serverless SQL pool. You need to monitor the number of queries that are currently executing. Which dynamic management view should you query?

A.sys.dm_resource_governor_workload_groups
B.sys.dm_exec_query_stats
C.sys.dm_exec_requests
D.sys.dm_exec_sessions
AnswerC

sys.dm_exec_requests shows currently executing requests in the serverless SQL pool, including state, command, and session ID.

Why this answer

Sys.dm_exec_returns detailed information about each request currently executing on the serverless SQL pool, including its state, command, and session ID. This DMV is specifically designed for monitoring active queries.

Option A (sys.dm_resource_governor_workload_groups) shows workload group configuration and resource statistics, not current requests.

Option B (sys.dm_exec_query_stats) provides cumulative performance statistics for cached query plans, not currently executing queries.

Option D (sys.dm_exec_sessions) contains session-level information but does not indicate which sessions are actively executing a request.

30
MCQmedium

You have an Azure Data Lake Storage Gen2 account used by an Azure Synapse Analytics serverless SQL pool. Analysts run ad-hoc queries against CSV and Parquet files. You need to reduce the amount of data scanned by these queries without changing file contents. What should you do?

A.Convert all CSV files to Parquet and store them in a single folder.
B.Create a partitioned folder hierarchy by date and query with OPENROWSET using a wildcard path.
C.Increase the DWU setting on the serverless SQL pool.
D.Enable hierarchical namespace on the storage account.
AnswerB

Partitioning folders by date lets the serverless SQL pool prune irrelevant directories, so a query filtered on date reads only the matching partition folders instead of the whole container. Combining that with an OPENROWSET wildcard path scoped to those folders directly reduces bytes scanned and cost, and it requires no rewrite of the underlying files.

Why this answer

Serverless SQL pools charge per byte scanned, so the effective lever is reducing the files a query must read. Organizing data into a folder hierarchy that mirrors common filter columns, such as date, allows partition elimination, and an OPENROWSET path that targets those folders avoids reading unrelated data. This keeps files unchanged while cutting scanned volume and cost.

Exam trap

The trap here is assuming that switching to a columnar format alone eliminates scanning, when without a filter-aligned folder layout the engine still reads every file in the container.

31
Multi-Selecteasy

Which TWO Azure services can be used to audit data access and changes in Azure Data Lake Storage Gen2? (Choose two.)

Select 2 answers
A.Microsoft Entra ID sign-in logs.
B.Azure Backup reports.
C.Storage account diagnostic settings.
D.Azure Monitor and Microsoft Sentinel.
E.Azure Policy.
AnswersC, D

Storage account diagnostic settings stream control-plane and data-plane events, including authenticated read, write and delete operations, to Log Analytics, storage or Event Hubs. This provides the audit trail of data access and changes required by the stem.

Why this answer

Storage account diagnostic settings (C) are correct because they can stream Data Lake Storage Gen2 resource logs—such as Read, Write, and Delete operations on blobs and ADLS Gen2 filesystems—to a Log Analytics workspace, storage account, or Event Hub, providing the raw audit trail of data access and changes. Azure Monitor and Microsoft Sentinel (D) are correct because Azure Monitor collects and queries those diagnostic logs (via Log Analytics/KQL) for auditing, while Microsoft Sentinel ingests the same storage/ADLS logs to detect, alert on, and investigate suspicious data-access activity. Microsoft Entra ID sign-in logs (A) only record authentication and token issuance events, not per-file data-plane operations, so they cannot audit data access or changes.

Azure Backup reports (B) cover backup job status and protected-item health, not data-plane read/write/delete auditing. Azure Policy (E) evaluates and enforces resource configuration compliance (control plane), not individual data access or modification events.

32
Multi-Selecthard

Which TWO Azure services can be used to monitor data pipeline runs and set up alerts for failures in Azure Data Factory?

Select 2 answers
A.Azure Data Factory monitoring views
B.Azure Log Analytics
C.Azure Sentinel
D.Azure Monitor
E.Azure Automation
AnswersB, D

Azure Log Analytics is part of Azure Monitor and is used to query pipeline logs and create alerts for failures.

Why this answer

Options B and D are correct. Azure Monitor is the primary service for collecting Azure Data Factory metrics and logs and for configuring alerts. Azure Log Analytics is used to query those logs and create failure alerts.

Option A is not a standalone Azure service; Azure Data Factory monitoring views are a built-in feature and do not independently provide alerting. Option C is a SIEM/SOAR security service, not for pipeline monitoring. Option E is for automation, not monitoring/alerting.

33
Multi-Selecteasy

Which TWO Azure features can be used to encrypt data at rest in Azure Blob Storage? (Choose two.)

Select 2 answers
A.Azure Disk Encryption
B.Azure Information Protection
C.Customer-managed keys in Azure Key Vault
D.Storage Service Encryption (SSE)
E.Transport Layer Security (TLS)
AnswersC, D

Customer-managed keys in Azure Key Vault satisfy the at-rest encryption requirement by letting you supply your own RSA key that wraps the account's data encryption key, rather than relying on Microsoft-managed keys. This gives you control over rotation and revocation, meeting the stem's demand for a Blob Storage encryption feature.

Why this answer

Option C (Customer-managed keys in Azure Key Vault) is correct because Azure Blob Storage supports server-side encryption with customer-managed keys (CMK), where the key encryption key is stored in Azure Key Vault and used to wrap the account's data encryption key, providing control over the encryption of data at rest. Option D (Storage Service Encryption, SSE) is correct because SSE is the built-in Azure Storage feature that automatically encrypts blob data at rest using AES-256 before it is persisted to disk, and it is enabled by default for all storage accounts. Option A (Azure Disk Encryption) is not correct because it encrypts OS and data disks attached to Azure VMs (using BitLocker/DM-Crypt), not blobs in a storage account.

Option B (Azure Information Protection) is not correct because it classifies and protects files at the application/document level rather than providing storage-account-level encryption at rest for blobs. Option E (Transport Layer Security) is not correct because TLS protects data in transit over the network, not data at rest.

34
MCQeasy

You need to ensure that an Azure Data Factory pipeline retries a failed activity up to three times with a 5-minute delay between retries. How should you configure the activity?

A.Configure the Retry policy on the pipeline activity as 'Exponential' with count 3
B.Set retry to 3 and retryIntervalInSeconds to 300 in the activity policy
C.Set the activity timeout to 15 minutes and enable retry
D.Set maxRetries to 3 and delay to 5 minutes in the pipeline JSON
AnswerB

The activity policy's retry property sets the maximum retry count, and retryIntervalInSeconds defines the pause between attempts. Setting retry to 3 and retryIntervalInSeconds to 300 yields three retries at five-minute intervals, exactly matching the stated requirement.

Why this answer

Azure Data Factory activity policies expose two retry-related properties: 'retry' (an integer count) and 'retryIntervalInSeconds' (the fixed delay between attempts). Setting retry to 3 and retryIntervalInSeconds to 300 produces exactly three retries with a 5-minute (300-second) fixed interval, matching the requirement. This is configured on the activity's policy object in the pipeline JSON or via the Author UI.

Exam trap

DP-203 often tests the exact ADF property names — candidates pick 'maxRetries' or 'Exponential' because those names appear in other Azure services, but ADF specifically uses 'retry' and 'retryIntervalInSeconds'.

How to eliminate wrong answers

Option A is wrong because ADF does not offer an 'Exponential' retry policy on activities — retry intervals in ADF are fixed, not exponential (that pattern exists in other services like AWS Step Functions or Azure Functions, not ADF activity policy). Option C is wrong because activity timeout controls how long a single execution may run before being cancelled; it does not control retry count or delay, and 15 minutes is unrelated to 3 retries × 5 minutes. Option D is wrong because 'maxRetries' and 'delay' are not the correct ADF property names — the actual schema uses 'retry' and 'retryIntervalInSeconds', so this JSON would fail validation.

35
Multi-Selecthard

Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to implement a monitoring strategy to detect and alert on unusual access patterns that could indicate a security breach. Which THREE services or features should you use? (Choose three.)

Select 3 answers
A.Enable Microsoft Defender for Storage to get security alerts about unusual access patterns.
B.Apply Azure Policy to enforce encryption and access policies.
C.Ingest the logs into Microsoft Sentinel and create analytics rules for anomalous patterns.
D.Enable diagnostic settings on the storage account to collect read, write, and delete logs.
E.Use Azure Monitor Metrics to track storage account transactions and latency.
AnswersA, C, D

Microsoft Defender for Storage continuously analyses data-plane telemetry and raises security alerts for anomalous access, such as unusual locations or suspicious enumeration. This directly satisfies the requirement to detect and alert on access patterns indicating a possible breach in the hierarchical-namespace account.

Why this answer

Option A is correct because Microsoft Defender for Storage analyzes data-plane telemetry on the storage account and raises security alerts for suspicious activity such as unusual access patterns, anomalous data exfiltration, or access from unusual locations, which directly addresses breach detection. Option C is correct because Microsoft Sentinel can ingest storage logs and use analytics rules (including built-in anomaly and threat-detection templates) to correlate events and alert on anomalous access patterns across the environment. Option D is correct because enabling diagnostic settings on the storage account is the prerequisite that exports read, write, and delete data-plane logs (the StorageRead/StorageWrite/StorageDelete categories) to a Log Analytics workspace, Event Hub, or storage account, providing the raw telemetry that Sentinel and Defender for Storage rely on for anomaly detection.

Option B does not belong because Azure Policy enforces configuration and compliance (for example, requiring encryption or HTTPS), but it does not detect or alert on unusual access patterns. Option E does not belong because Azure Monitor Metrics provides aggregate numeric time-series data such as transaction counts, latency, and availability, not per-request access detail needed to identify anomalous access patterns.

36
MCQeasy

You use Azure Data Lake Storage Gen2 with a hierarchical namespace. You need to delegate permissions to a group of data scientists so they can create folders and upload files only within a specific directory path. What is the best way to achieve this?

A.Use a stored access policy to grant permissions to the directory.
B.Set ACL entries on the specific directory path granting read, write, and execute permissions to the users.
C.Generate a shared access signature (SAS) with permissions scoped to the specific directory.
D.Assign the Storage Blob Data Contributor role to the users at the storage account level.
AnswerB

POSIX ACLs on the target directory grant read, write and execute to the scientists, scoping access to that path only. This satisfies the constraint of limiting folder creation and uploads to a specific directory, which RBAC roles cannot do at path level.

Why this answer

Azure Data Lake Storage Gen2 with hierarchical namespace supports POSIX-like ACLs at the directory and file level. To delegate permissions to a group of data scientists so they can create folders and upload files only within a specific directory path, you set ACL entries on that directory granting the necessary permissions (read, write, execute) to the group. This allows granular access control without granting broader permissions at the storage account level.

Exam trap

The trap here is confusing SAS tokens with ACLs. Many candidates think SAS is the go-to for granular access, but in ADLS Gen2 with hierarchical namespace, ACLs are the native and recommended method for directory-level permissions. Also, some might choose the broad role assignment for simplicity, ignoring the least privilege requirement.

How to eliminate wrong answers

Option A is wrong because a stored access policy is used for shared access signatures (SAS) on containers, not for direct ACL management on directories. Option C is wrong because a SAS provides temporary, token-based access and is not the best method for persistent, directory-level permission delegation to a group; it also lacks the granularity of ACLs for hierarchical namespaces. Option D is wrong because assigning the Storage Blob Data Contributor role at the storage account level grants overly broad permissions across the entire account, violating the principle of least privilege.

37
MCQeasy

You are configuring security for an Azure Data Lake Storage Gen2 account that stores sensitive data. You need to ensure that all data access is logged and that you can audit who accessed the data and when. You also need to retain the logs for 90 days. What should you do?

A.Configure diagnostic settings to send logs to a Log Analytics workspace.
B.Enable soft delete for blobs.
C.Use Azure Storage Analytics logging to a storage account.
D.Enable Azure Defender for Storage.
AnswerA

Configuring diagnostic settings for Azure Data Lake Storage Gen2 allows you to stream resource logs, such as StorageRead and StorageWrite, to a Log Analytics workspace. These logs capture details about each access, including the identity, operation, and timestamp. You can then query and retain the logs in Log Analytics for 90 days (or more) to meet auditing requirements.

Why this answer

Diagnostic settings in Azure allow you to export resource logs to a Log Analytics workspace, where you can analyze and retain them for up to 90 days (or longer with custom retention). This provides the necessary auditing capability for Data Lake Storage Gen2 access.

Exam trap

The trap here is confusing threat detection with access logging; Azure Defender alerts on suspicious activity but does not provide a full audit trail.

38
MCQeasy

Your company uses Azure Data Lake Storage Gen2. You need to ensure that data at rest is encrypted using a customer-managed key stored in Azure Key Vault. What should you configure?

A.Use Azure Policy to audit storage accounts without encryption.
B.Enable 'Azure Storage encryption' with customer-managed keys in the storage account's encryption blade.
C.Implement client-side encryption in the application code.
D.Enable 'Infrastructure encryption' for double encryption.
AnswerB

Configuring customer-managed keys in the storage account's encryption blade satisfies the requirement for encryption at rest using a key held in Azure Key Vault. Azure Storage encryption applies AES-256 to all data at rest, and selecting customer-managed keys delegates key control to Key Vault rather than Microsoft-managed keys.

Why this answer

Azure Storage encryption with customer-managed keys is configured in the encryption blade of the storage account. This ensures data at rest is encrypted using a key stored in Azure Key Vault. Option A is incorrect because Azure Policy can audit or enforce encryption but does not configure customer-managed keys.

Option C is incorrect because client-side encryption encrypts data before it reaches Azure Storage, not at rest. Option D is incorrect because infrastructure encryption provides a second encryption layer but does not use customer-managed keys for the primary encryption.

39
MCQmedium

You manage an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline runs daily and completes successfully. You need to be alerted when the pipeline duration exceeds 60 minutes. You want to minimize administrative effort. What should you do?

A.Add a Web activity in the pipeline that calls a Logic App to send an email if the pipeline runs longer than 60 minutes.
B.Configure a diagnostic setting to send activity logs to a Log Analytics workspace, then create a log search alert rule using a query that filters for pipeline runs with duration greater than 60 minutes.
C.Create an Azure Monitor alert rule on the ADFPipelineRun metric with a threshold of 60 minutes.
D.Enable Azure Monitor alerts on the Failed Runs metric of the pipeline and set the threshold to 1.
AnswerB

Diagnostic settings route Azure Data Factory activity logs, including pipeline run records with a Duration property, to Log Analytics. A log search alert rule can run a scheduled query such as ADFPipelineRun | where Duration > 60m and trigger an action group when results are found. This is the standard, low-effort method to alert on pipeline duration without custom code or external monitoring.

Why this answer

To alert on pipeline duration, you need access to run-level logs that include the Duration property. Diagnostic settings export Azure Data Factory activity logs to Log Analytics, where you can write a query to find runs exceeding 60 minutes and attach an alert rule. This is the least-effort, native Azure solution.

Metric-based alerts on pipeline metrics like ADFPipelineRun or Failed Runs do not capture duration, and embedding custom activities adds unnecessary complexity.

Exam trap

The trap here is assuming that Azure Monitor metric alerts can directly monitor pipeline duration, when duration is only available in activity logs requiring a log search alert.

40
MCQeasy

You are responsible for securing an Azure Synapse Analytics workspace that contains sensitive data. You need to ensure that data is encrypted at rest using a customer-managed key stored in Azure Key Vault. What should you configure?

A.Always Encrypted with column encryption keys.
B.Transparent Data Encryption (TDE) with a customer-managed key.
C.Transparent Data Encryption (TDE) with a service-managed key.
D.Azure Storage Service Encryption with a customer-managed key.
AnswerB

TDE with a customer-managed key allows you to bring your own key stored in Azure Key Vault. This gives you control over the encryption key, including rotation and revocation. It satisfies the requirement to encrypt data at rest using a customer-managed key in Azure Key Vault.

Why this answer

Transparent Data Encryption (TDE) with a customer-managed key in Azure Key Vault provides encryption at rest for Azure Synapse Analytics dedicated SQL pools. This allows you to manage the encryption key, including rotation and revocation, and meets compliance requirements for customer-managed keys.

Exam trap

The trap here is confusing Always Encrypted with TDE, or assuming that Azure Storage Service Encryption applies to Synapse dedicated SQL pools.

41
MCQhard

You are designing a data lake in Azure Data Lake Storage Gen2 for a large enterprise. You need to ensure that only authorized users can access the data, and you must implement the principle of least privilege. Which security mechanism should you use to grant fine-grained access to specific directories and files without modifying the underlying storage account firewall settings?

A.Azure RBAC roles combined with POSIX-like ACLs
B.Managed identities for Azure resources
C.Storage account firewall rules
D.Shared access signatures (SAS)
AnswerA

Azure RBAC grants coarse access at account, container or folder scope, while POSIX-like ACLs refine permissions down to individual directories and files for specific principals. Combining both enforces least privilege without touching the storage account firewall.

Why this answer

Azure RBAC roles combined with POSIX-like ACLs is correct because Azure Data Lake Storage Gen2 supports a hierarchical namespace that enables POSIX-style access control lists (ACLs) at the directory and file level, allowing fine-grained permissions (read, write, execute) for specific Azure AD principals. RBAC handles coarse-grained control at the management group, subscription, or storage account scope, while ACLs provide the granular, object-level authorization required for least privilege. This combination does not require altering the storage account firewall, which controls network access rather than identity-based permissions.

Exam trap

DP-203 often tests the misconception that storage account firewall rules or SAS tokens provide identity-based, fine-grained access control, when in fact only the combination of RBAC and POSIX ACLs delivers directory-level least privilege without network changes.

How to eliminate wrong answers

Option B is wrong because managed identities are used to authenticate Azure resources to other services, not to grant fine-grained access to directories and files for users. Option C is wrong because storage account firewall rules restrict network access based on IP address or virtual network, not user identity or directory-level permissions. Option D is wrong because shared access signatures (SAS) grant time-limited, delegated access via a token but do not provide the persistent, identity-based, fine-grained ACL model needed for least privilege across many users and directories.

42
MCQeasy

You have an Azure Data Lake Storage Gen2 account that contains sensitive data. You need to ensure that data is encrypted at rest and that you control the encryption keys. You also need to be able to audit key usage. What should you implement?

A.Azure Information Protection with Bring Your Own Key (BYOK).
B.Customer-managed keys (CMK) in Azure Key Vault for Azure Storage encryption.
C.Storage Service Encryption with Microsoft-managed keys and Azure Monitor logs.
D.Azure Disk Encryption with Azure Key Vault.
AnswerB

Customer-managed keys allow you to use your own keys in Azure Key Vault to encrypt the storage account. This gives you control over key rotation and access, and you can audit key usage through Key Vault logging. It meets the requirements for encryption at rest, key control, and auditing.

Why this answer

Customer-managed keys in Azure Key Vault allow you to encrypt an Azure Storage account with your own keys, providing control over key lifecycle and enabling auditing via Key Vault logs. This satisfies the need for encryption at rest, key control, and auditability. Other options either do not apply to storage accounts or do not grant key control.

Exam trap

The trap here is confusing Azure Disk Encryption with Storage Service Encryption. Azure Disk Encryption encrypts VM disks, not storage accounts, and does not provide the required control over keys for Data Lake Storage Gen2.

43
MCQmedium

Your Azure Synapse Analytics dedicated SQL pool is experiencing performance degradation. You notice that some queries are being queued due to resource class conflicts. What should you implement to optimize performance and reduce queuing?

A.Scale the dedicated SQL pool to a higher DWU level
B.Configure workload management with workload groups and classifiers
C.Create materialized views for the most common aggregations
D.Enable result-set caching for frequently run queries
AnswerB

Workload groups with classifiers assign queries to resource buckets based on rules, giving each workload dedicated memory and concurrency. This reduces resource class conflicts and queuing by preventing competing queries from contending for the same resources.

Why this answer

Workload management with workload groups and classifiers allows you to assign queries to different resource classes and prioritize them, directly addressing resource class conflicts and reducing queuing. Option A is incorrect: scaling the pool to a higher DWU increases overall resources but does not specifically manage resource class contention; it may also incur additional cost without solving the root issue. Option C is incorrect: materialized views improve query performance by pre-aggregating data but do not affect concurrency or queuing.

Option D is incorrect: result-set caching reduces repeated computation for identical queries but does not resolve queuing caused by resource class conflicts.

44
MCQeasy

You have an Azure Data Factory pipeline that copies data from an FTP server to Azure Blob Storage. The pipeline runs successfully most of the time, but occasionally fails with a 'FTP server connection refused' error during peak hours. You need to minimize these failures with minimal cost. What should you do?

A.Add a retry policy to the copy activity with a backoff interval.
B.Set up Azure ExpressRoute to improve network reliability.
C.Migrate the FTP server to SFTP.
D.Increase the parallel copy count in the copy activity.
AnswerA

Transient connection refusals during peak load are best absorbed by retry logic. A retry policy with backoff interval reattempts the copy activity after progressively longer waits, resolving intermittent FTP failures at negligible cost and no infrastructure changes.

Why this answer

A retry policy with a backoff interval on the copy activity handles transient 'connection refused' errors during peak hours by automatically reattempting after a delay, smoothing over temporary FTP server saturation. It is the lowest-cost, least-invasive fix and directly targets intermittent failures without infrastructure changes.

Exam trap

The trap is choosing an infrastructure or protocol change (ExpressRoute, SFTP) for what is a transient, intermittent error — the exam rewards the minimal-cost, built-in retry mechanism.

How to eliminate wrong answers

Option B is wrong because ExpressRoute is a dedicated private network connection with significant cost and provisioning time, and it does not address FTP server-side connection limits during peak load. Option C is wrong because migrating to SFTP changes the protocol and security posture but does not resolve connection-refused errors caused by server capacity or connection limits. Option D is wrong because increasing parallel copy count would open more concurrent connections, likely worsening the FTP server's connection-refused condition rather than alleviating it.

45
MCQhard

You are designing a data processing solution using Azure Synapse Analytics serverless SQL pool. The solution will query data stored in Parquet files in Azure Data Lake Storage Gen2. You need to ensure that the queries are optimized for performance. Which action should you take?

A.Increase the MAXDOP setting in the query.
B.Convert the Parquet files to CSV format for faster parsing.
C.Create materialized views on the external tables.
D.Partition the Parquet files by date and use partition pruning in the query.
AnswerD

Partitioning Parquet files by date lets the serverless SQL pool read only relevant folders, since partition pruning eliminates scanning of unrelated files. This directly reduces data scanned and I/O, which is the dominant performance constraint for serverless queries over Data Lake Storage Gen2.

Why this answer

In Azure Synapse serverless SQL pool, performance is dominated by how much data must be read from storage, since there is no persistent compute or indexing. Partitioning Parquet files by date and filtering on the partition column enables partition pruning, so the engine reads only the relevant folders instead of scanning the entire dataset. This dramatically reduces I/O and query time, making D the correct optimization.

Exam trap

DP-203 often tests whether candidates know serverless SQL pool limitations; the trap is picking materialized views (a dedicated pool feature) or MAXDOP tuning, which do not address the I/O-bound nature of serverless queries.

How to eliminate wrong answers

Option A is wrong because MAXDOP controls the degree of parallelism for a query, but serverless SQL pool manages parallelism automatically and MAXDOP does not reduce the volume of data scanned — the primary bottleneck. Option B is wrong because CSV is row-based and uncompressible compared to columnar Parquet; converting to CSV would increase I/O and slow parsing, the opposite of optimization. Option C is wrong because materialized views are not supported on external tables in serverless SQL pool (they are a dedicated SQL pool feature), so this action is not even possible.

46
MCQmedium

You are a data engineer at a healthcare company. Your Azure Synapse Analytics workspace contains a dedicated SQL pool that holds patient records. A new compliance rule requires that all queries against the dedicated SQL pool be audited, and that any attempt to access data from an unauthorized IP address be logged. You need to configure auditing for the dedicated SQL pool. What should you do?

A.Enable auditing on the Azure Synapse Analytics workspace and set the audit log destination to a Log Analytics workspace.
B.Enable Microsoft Defender for SQL on the dedicated SQL pool and set the audit log destination to Azure Storage.
C.Configure Azure SQL Database auditing for the dedicated SQL pool by using the master database.
D.Enable auditing on the dedicated SQL pool and configure the audit log destination to Azure Storage, Log Analytics, or Event Hubs.
AnswerD

Dedicated SQL pool auditing is configured directly on the pool and supports Azure Storage, Log Analytics, and Event Hubs as destinations. This captures all queries and unauthorized IP access attempts, meeting the compliance requirement. The setting is found in the Azure portal under the dedicated SQL pool's Security section, and it can also be set via PowerShell or Azure CLI.

Why this answer

Auditing for an Azure Synapse Analytics dedicated SQL pool must be enabled directly on the pool, not at the workspace or Azure SQL Database level. The audit logs can be sent to Azure Storage, Log Analytics, or Event Hubs, providing a full record of queries and unauthorized access attempts. This configuration satisfies the compliance requirement to audit all queries and log unauthorized IP access attempts.

Exam trap

The trap here is assuming that workspace-level auditing or Microsoft Defender for SQL automatically covers dedicated SQL pool auditing, when in fact dedicated SQL pool auditing is configured separately on the pool itself.

47
MCQmedium

You are a data engineer at a healthcare company. You have an Azure Data Lake Storage Gen2 account named sthealthcare with a container named records. The container holds sensitive patient data in Parquet files. You need to ensure that only users who are members of the Azure AD group named ClinicalResearchers can read the data, while users in the group DataEngineers can read and write. Access must be managed at the directory level and must not affect other containers in the storage account. What should you do?

A.Assign the Storage Blob Data Reader role to ClinicalResearchers and Storage Blob Data Contributor role to DataEngineers at the storage account scope.
B.Enable hierarchical namespace on the storage account, then set POSIX-like ACLs on the records directory, granting read to ClinicalResearchers and read-write to DataEngineers.
C.Configure Azure RBAC roles on the records container: assign Storage Blob Data Reader to ClinicalResearchers and Storage Blob Data Contributor to DataEngineers.
D.Create a shared access signature (SAS) token for each user in ClinicalResearchers and DataEngineers with appropriate permissions, and distribute the tokens.
AnswerB

Azure Data Lake Storage Gen2 supports POSIX-like ACLs when hierarchical namespace is enabled. Applying ACLs at the directory level allows granular control per user or group without affecting other containers. This meets the requirement for directory-level access and least privilege. The groups must exist in Azure AD and be assigned as principals on the ACL entries.

Why this answer

The requirement is to manage access at the directory level within a container without affecting other containers. Azure Data Lake Storage Gen2 ACLs, enabled by hierarchical namespace, allow setting permissions on directories and files. By applying ACLs on the records directory, you grant the appropriate groups read or read-write access only to that directory and its contents.

This satisfies least privilege and directory-level control.

Exam trap

The trap here is assuming that Azure RBAC role assignments at the container level provide directory-level granularity, when in fact only POSIX ACLs on a hierarchical namespace-enabled account can enforce per-directory permissions.

48
Multi-Selecthard

Your organization uses Azure Purview for data governance. You need to ensure that sensitive data is properly classified and that access to it is monitored. Which THREE actions should you take? (Choose three.)

Select 3 answers
A.Define Azure Policy initiatives to enforce classification on all storage accounts.
B.Use Azure Sentinel to classify data as it is ingested.
C.Create custom sensitivity labels in Microsoft Purview Information Protection and apply them to data sources.
D.Integrate Azure Purview with Microsoft Defender for Cloud Apps to monitor access to sensitive data.
E.Set up automated scanning in Azure Purview to discover and classify sensitive data.
AnswersC, D, E

Custom sensitivity labels in Microsoft Purview Information Protection extend classification beyond built-in types, letting you tag organisation-specific sensitive data and apply those labels across registered data sources. This satisfies the stem's requirement that sensitive data be properly classified, while label usage feeds Purview's monitoring of access to labelled assets.

Why this answer

Option C is correct because Microsoft Purview Information Protection sensitivity labels (created in the compliance portal) are the mechanism for tagging data with sensitivity classifications, and Purview can extend these labels to data sources such as Azure SQL, Storage, and Synapse for consistent classification. Option D is correct because integrating Azure Purview with Microsoft Defender for Cloud Apps enables monitoring of user access and activities on sensitive data that Purview has classified, providing anomaly detection and governance over access. Option E is correct because Azure Purview's automated scanning uses built-in and custom classification rules to discover and classify sensitive data (for example, credit card or national ID patterns) across registered sources, which is the core way to ensure data is properly classified.

Option A is not correct because Azure Policy initiatives enforce configuration and compliance states on resources; they do not perform data classification of the contents within storage accounts. Option B is not correct because Azure Sentinel is a SIEM/SOAR solution for security analytics and threat detection, not a data classification engine for ingested data.

Exam trap

DP-203 often tests the boundary between governance tools — candidates confuse Azure Policy (resource configuration enforcement) and Azure Sentinel (SIEM) with Purview's actual classification and monitoring capabilities, picking them as if they performed data classification.

49
MCQeasy

You need to monitor the performance of an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Blob Storage. The pipeline runs on a self-hosted integration runtime. Which metric is most important to monitor to ensure the self-hosted IR is not a bottleneck?

A.Pipeline duration metric
B.Queue depth for the self-hosted IR
C.Number of active connections to the IR
D.Data read and data written metrics for the pipeline
AnswerB

Queue depth measures pending requests awaiting a free self-hosted integration runtime node, so a rising value directly signals the IR cannot keep pace with copy activity demand. This exposes the bottleneck constraint, whereas CPU or memory alone may not reflect queued work.

Why this answer

The self-hosted integration runtime (IR) queue depth metric indicates the number of activities waiting to be processed by the IR. A consistently high or increasing queue depth signals that the IR cannot keep up with the workload, making it a bottleneck. Monitoring this metric allows proactive scaling or optimization.

Exam trap

DP-203 often tests the distinction between pipeline-level metrics (duration, data read/written) and IR-specific metrics (queue depth, CPU); candidates may pick duration as a proxy for IR performance.

How to eliminate wrong answers

Option A is wrong because pipeline duration is an end-to-end metric that can be affected by many factors (source, sink, network), not specifically the IR. Option C is wrong because active connections do not directly indicate IR saturation; the IR can handle many connections if not CPU/memory bound. Option D is wrong because data read/written metrics reflect throughput but not whether the IR is the limiting factor; they are outcomes, not bottleneck indicators.

50
MCQeasy

An organization is using Azure Synapse Analytics and wants to implement column-level security to restrict access to sensitive columns. Which feature should they use?

A.Dynamic data masking
B.Azure Purview
C.Column-level security using GRANT
D.Row-level security
AnswerC

Column-level security in Azure Synapse Analytics is enforced through T-SQL GRANT statements, letting you deny SELECT on specific columns while permitting access to the rest of the table. This directly satisfies the requirement to restrict sensitive columns, unlike object-level permissions or dynamic data masking, which obscures values rather than blocking access.

Why this answer

Column-level security in Azure Synapse Analytics is implemented using GRANT statements on specific columns, restricting access to sensitive columns. Option A is incorrect because dynamic data masking obfuscates data at query time but does not prevent access. Option B is incorrect because Azure Purview is a data governance service, not for access control.

Option D is incorrect because row-level security filters rows, not columns.

51
MCQhard

You have an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline uses a self-hosted integration runtime. You need to ensure that data is encrypted in transit and that the integration runtime authenticates to the on-premises SQL Server using Windows authentication. What should you configure?

A.Use Azure ExpressRoute to connect to the on-premises network and configure the linked service to use SQL authentication.
B.Set up a VPN gateway between Azure and the on-premises network and use a service principal for authentication.
C.Configure the self-hosted integration runtime to use a managed identity and enable Always Encrypted on the SQL Server.
D.Enable SSL encryption on the SQL Server connection and configure the linked service to use Windows authentication with a domain account.
AnswerD

To encrypt data in transit between the self-hosted integration runtime and the on-premises SQL Server, you must enable SSL/TLS on the SQL Server connection by setting Encrypt=True in the connection string. For Windows authentication, the linked service must use a domain account with the necessary permissions. This combination ensures both secure transmission and proper authentication, meeting the requirements.

Why this answer

For a self-hosted integration runtime copying from on-premises SQL Server, data in transit is encrypted by enabling SSL on the SQL Server connection (Encrypt=True). Windows authentication requires a domain account configured in the linked service. Combining these ensures secure transmission and proper authentication.

The other options either do not provide the necessary encryption, use incompatible authentication methods, or address network-level security instead of the application-level requirements.

Exam trap

The trap here is confusing network-level encryption (VPN, ExpressRoute) with application-level encryption (SSL/TLS) and assuming that Azure AD authentication works for on-premises SQL Server.

52
MCQmedium

You manage an Azure Synapse Analytics workspace. A dedicated SQL pool contains a table with a column named CustomerEmail that stores email addresses. You need to ensure that users who are not members of the DataPrivacy role see only a masked version of the email addresses when they query the table, while members of DataPrivacy see the actual values. The solution must minimize administrative effort. What should you do?

A.Enable Always Encrypted with a column encryption key and configure the DataPrivacy role as the only role with access to the key.
B.Create a dynamic data masking rule on the CustomerEmail column using the email() masking function and grant UNMASK to the DataPrivacy role.
C.Apply a column-level security policy using GRANT SELECT ON the CustomerEmail column to the DataPrivacy role.
D.Implement row-level security with a predicate function that filters rows based on the user's role.
AnswerB

Dynamic data masking applies a masking rule to the column so that non-privileged users see masked data, while users with the UNMASK permission see the actual values. This directly meets the requirement with minimal administrative effort because the masking rule and UNMASK grant are set once and apply to all queries.

Why this answer

Dynamic data masking is designed to limit sensitive data exposure by masking it to non-privileged users. Applying an email() mask to the CustomerEmail column ensures that users without the UNMASK permission see a masked format, while granting UNMASK to the DataPrivacy role allows those members to see the actual values. This approach requires only a one-time configuration and no changes to queries or applications.

Exam trap

The trap here is confusing dynamic data masking with column-level security or encryption, which control access rather than present masked data to unauthorized users.

53
MCQmedium

You are reviewing a script to create an external data source in Azure Synapse Analytics serverless SQL pool. Based on the exhibit, what is the purpose of the SAS token?

A.To provide read access to the container for querying data.
B.To provide write access to the container for storing query results.
C.To encrypt the connection between the serverless pool and storage.
D.To authenticate the user to the serverless SQL pool.
AnswerA

The SAS token supplies delegated, time-limited credentials appended to the external data source location, letting the serverless SQL pool authenticate to the storage container. Without it, queries against the external data source would fail authorisation, so it grants the read access needed for querying.

Why this answer

In Azure Synapse Analytics serverless SQL pool, an external data source pointing to Azure Storage requires a SAS token to provide read access to the container so that queries can read the data. The SAS token grants limited access rights without exposing the account key, and it is used in the CREATE EXTERNAL DATA SOURCE statement to authenticate and authorize access to the storage.

Exam trap

DP-203 often tests the purpose of SAS tokens in external data sources; candidates might confuse SAS tokens with account keys or think they provide write access, but they are primarily for delegated read access.

How to eliminate wrong answers

Option B is wrong because the SAS token is used for read access to query data, not for write access to store query results; serverless SQL pool query results are typically stored elsewhere or returned to the client. Option C is wrong because encryption of the connection is handled by HTTPS, not by the SAS token; the SAS token is for authorization. Option D is wrong because authentication to the serverless SQL pool is handled by Azure AD or SQL authentication, not by a SAS token; the SAS token authenticates to the storage account.

54
MCQeasy

You need to monitor the performance of an Azure Stream Analytics job that processes real-time IoT data. Which metric indicates the number of events that are being dropped or delayed due to insufficient processing capacity?

A.Watermark delay.
B.Output events.
C.Backlogged input events.
D.Input events.
AnswerC

Backlogged input events counts events queued but not yet processed because the job lacks sufficient streaming units. Rising values indicate the job cannot keep pace with incoming IoT data, signalling that events are delayed or dropped due to inadequate processing capacity.

Why this answer

Backlogged input events (C) measures the number of events that are queued awaiting processing, indicating that the job is unable to keep up with the input rate. High backlog suggests insufficient processing capacity, leading to dropped or delayed events. Watermark delay (A) measures the time lag in processing, but does not directly count events dropped.

Input events (D) is the total received, not dropped/delayed. Output events (B) is the total sent.

55
Multi-Selecthard

Which THREE metrics should you monitor to evaluate the performance of an Azure Stream Analytics job?

Select 3 answers
A.Input Events Backlogged
B.Output Events
C.Conversion Errors
D.SU (Memory) Utilization
E.Watermark Delay (seconds)
AnswersA, B, E

Input Events Backlogged counts events awaiting processing, exposing whether the job's streaming units or query logic cannot keep pace with incoming throughput. Rising backlog directly signals a performance bottleneck in the Azure Stream Analytics pipeline.

Why this answer

Input Events Backlogged (A) is a correct metric because it measures the number of input events that are waiting to be processed, directly revealing whether the job is falling behind its input rate. Output Events (B) is correct because it counts the events emitted to the output sink, letting you verify the job is actually producing results and detect drops in throughput. Watermark Delay (seconds) (E) is correct because it quantifies how far behind real time the job is running, which is the key indicator of streaming latency and timeliness.

Conversion Errors (C) is not one of the three performance metrics for evaluating job throughput/latency; it reflects data-format deserialization problems rather than performance. SU (Memory) Utilization (D) is a Streaming Unit resource metric that indicates capacity consumption, not the job's runtime performance as measured by backlog, output, and watermark delay.

56
MCQmedium

You are designing a data processing solution in Azure Synapse Analytics. The solution must ensure that sensitive columns containing personally identifiable information (PII) are masked at query time for users without explicit permissions. Which Azure Synapse Analytics feature should you use?

A.Row-Level Security
B.Transparent Data Encryption
C.Dynamic Data Masking
D.Always Encrypted
AnswerC

Dynamic Data Masking applies masking rules to designated columns so unauthorised users see obfuscated values at query time, while privileged users retain full data. This satisfies the requirement to mask PII without altering stored data.

Why this answer

Dynamic Data Masking (DDM) in Azure Synapse Analytics masks sensitive column data at query time for users without explicit UNMASK permissions, returning masked values (e.g., XXXX or 0) instead of the real data. It is applied via a MASKED WITH clause on the column definition and does not alter the stored data, making it the correct choice for query-time PII obfuscation.

Exam trap

The trap here is confusing encryption-at-rest (TDE, Always Encrypted) with query-time masking; candidates often pick Always Encrypted thinking it hides data from unauthorized users, but it actually requires the client to decrypt, so it does not mask for users without permissions.

How to eliminate wrong answers

Option A is wrong because Row-Level Security filters which rows a user can see based on a predicate, not which column values are masked. Option B is wrong because Transparent Data Encryption encrypts data at rest on disk and is transparent to all authorized queries — it does not mask values for specific users. Option D is wrong because Always Encrypted encrypts data client-side and requires the client to hold the column master key; it protects data in transit and at rest but does not provide query-time masking for users lacking permissions.

57
MCQeasy

You are designing a security strategy for an Azure Data Lake Storage Gen2 account that stores sensitive data. You need to ensure that data is encrypted at rest using customer-managed keys. What should you configure?

A.Enable Azure Storage Service Encryption with Microsoft-managed keys.
B.Enable infrastructure encryption (double encryption) on the storage account.
C.Use Azure Disk Encryption on the virtual machines that access the data lake.
D.Configure a customer-managed key in Azure Key Vault and associate it with the storage account.
AnswerD

To use customer-managed keys for Azure Storage encryption, you create a key in Azure Key Vault (or Managed HSM) and then configure the storage account to use that key for encryption at rest. This allows you to control key lifecycle, rotation, and access. The storage account must have a managed identity with permissions to access the key vault. This is the standard method to meet the requirement of customer-managed keys for data at rest.

Why this answer

Customer-managed keys for Azure Storage encryption require you to create and manage a key in Azure Key Vault and then configure the storage account to use that key. This gives you control over encryption at rest. Microsoft-managed keys are the default but do not meet the requirement.

Infrastructure encryption adds a second layer but does not change key management. Azure Disk Encryption protects VM disks, not the data lake. Therefore, the correct approach is to use a customer-managed key from Key Vault.

Exam trap

The trap here is confusing default storage encryption (Microsoft-managed keys) or infrastructure encryption with the specific configuration needed for customer-managed keys.

58
MCQeasy

Your team has deployed an Azure Stream Analytics job that writes output to Azure Cosmos DB. You need to monitor the job for data latency and ensure it meets a service-level agreement (SLA) of under 10 seconds from input to output. Which metric should you track in Azure Monitor?

A.Output events.
B.Runtime errors.
C.Watermark delay.
D.Input events.
AnswerC

Watermark delay measures the difference between the latest event processed and the newest event received, expressed in seconds. It directly quantifies end-to-end processing lag, so comparing it against the 10-second SLA threshold tells you whether the job meets the required latency.

Why this answer

Watermark delay is the correct metric to monitor for data latency because it measures the maximum time between an input event being received and the corresponding output being produced. A watermark delay consistently under 10 seconds ensures the SLA is met. Output events (A) track the number of output events, not latency.

Runtime errors (B) indicate failures, not latency. Input events (D) track the number of input events, not latency.

59
MCQeasy

Your organization uses Microsoft Purview to catalog data assets. You need to ensure that sensitive data such as credit card numbers are automatically detected and labeled. Which Purview feature should you configure?

A.Create an Azure Policy to enforce tagging.
B.Configure a scan rule set with built-in classification rules for sensitive data types.
C.Enable the Data Catalog self-service search.
D.Enable Microsoft Information Protection for the data sources.
AnswerB

Scan rule sets bundle classification rules, including built-in system rules for credit card and other sensitive types, which Purview applies during scans to detect and label matching data automatically. This satisfies the automatic detection requirement in the stem.

Why this answer

Microsoft Purview automatically detects sensitive data types such as credit card numbers through classification rules that are part of a scan rule set. When you configure a scan rule set and include the built-in system classification rules (e.g., Credit Card Number, which matches patterns like 16-digit numbers with Luhn validation), Purview applies those classifications during scans and can then apply sensitivity labels. This is the native mechanism for automated sensitive data detection and labeling in Purview.

Exam trap

The trap is confusing classification with labeling or with policy enforcement; Purview scans and classifies sensitive data, but labels are applied through auto-labeling policies or MIP, not by Azure Policy or the catalog search.

How to eliminate wrong answers

Option A is wrong because Azure Policy enforces resource tagging and compliance at the Azure control plane; it does not scan data contents or classify sensitive data types like credit card numbers. Option C is wrong because Data Catalog self-service search is a discovery/browsing feature for users, not a detection or labeling mechanism. Option D is wrong because Microsoft Information Protection (MIP) provides labeling and protection capabilities but does not itself perform the automated content scanning and classification of data sources; that is done by Purview's scanning and classification engine, which can then integrate with MIP labels.

60
MCQhard

Your organization uses Azure Data Lake Storage Gen2 with hierarchical namespace enabled. You need to grant a service principal read and write access to a specific directory without granting access to the parent directories. What should you use?

A.Assign the Storage Blob Data Contributor role at the directory level using RBAC.
B.Use a managed identity and assign it to the directory.
C.Create a stored access policy on the directory.
D.Set ACLs on the directory with default ACLs for the service principal.
AnswerD

Default ACLs apply only to new child items created within the directory, not to the directory itself or to existing objects, so they cannot grant the service principal read and write access to the target directory directly. This option is tempting because default ACLs are designed to propagate permissions to future files and subdirectories, making them correct when the requirement is to control access for newly created content rather than the existing directory.

Why this answer

Azure RBAC roles for Azure Storage cannot be scoped to a directory or file; they are assigned at the storage account or container level. To grant a service principal read and write access to a specific directory in an Azure Data Lake Storage Gen2 account with hierarchical namespace enabled, use access control lists (ACLs) on that directory. Default ACLs on the directory will be inherited by newly created child items.

Option A is incorrect because the Storage Blob Data Contributor role cannot be assigned at the directory level. Option B is incorrect because a managed identity is an identity, not a permission mechanism. Option C is incorrect because stored access policies are used for shared access signatures (SAS), not for granting directory access.

61
Multi-Selectmedium

You are monitoring an Azure Data Factory pipeline that copies data from an on-premises SQL Server to Azure Data Lake Storage Gen2. The pipeline runs daily and has recently started taking longer than expected. You need to identify the cause of the performance degradation. Which two actions should you perform? (Choose two.)

Select 2 answers
A.Increase the DIU (Data Integration Units) for the copy activity to improve throughput.
B.Check the on-premises SQL Server performance counters for CPU, memory, and disk I/O during the pipeline run.
C.Review the Data Factory activity run logs in Azure Monitor to identify the duration of each activity.
D.Reconfigure the pipeline to use a self-hosted integration runtime with more nodes.
E.Enable Azure Data Factory diagnostic settings to send logs to Log Analytics and analyze with Kusto queries.
AnswersB, C

The source SQL Server could be the bottleneck due to resource contention. Monitoring its performance counters during the pipeline run can reveal if the server is under heavy load, causing slower data reads. This helps identify whether the degradation is due to source-side issues, which is a common cause in on-premises to cloud copy scenarios.

Why this answer

Reviewing activity run logs helps identify which activity is slow and its duration, while checking on-premises SQL Server performance counters determines if the source is a bottleneck. Together, these diagnostic steps pinpoint the cause of the performance degradation. The other options are either remediation actions or broader monitoring setups that do not directly diagnose the recent slowdown.

Exam trap

The trap here is jumping to remediation like increasing DIU or scaling integration runtime without first diagnosing the root cause of the performance degradation.

62
MCQeasy

You are monitoring an Azure Data Factory pipeline that runs hourly. The pipeline executes a stored procedure in an Azure SQL Database. Recently, you have observed that the pipeline occasionally fails with a 'Deadlock' error when the stored procedure runs. The Azure SQL Database is configured with the 'Read Committed Snapshot' isolation level enabled. You need to resolve the deadlock issue with minimal impact on performance. The stored procedure updates multiple tables in a single transaction and is critical for reporting. What should you do?

A.Change the stored procedure to use NOLOCK hints
B.Remove the transaction from the stored procedure
C.Add retry logic in the Data Factory pipeline for the stored procedure activity
D.Disable the 'Read Committed Snapshot' isolation level
AnswerC

Retry logic in the Data Factory activity handles transient deadlock failures without altering the stored procedure's transaction semantics or isolation level. Since the procedure updates multiple tables in one transaction and is critical for reporting, retrying the activity preserves Read Committed Snapshot behaviour and avoids performance penalties, satisfying the minimal-impact constraint.

Why this answer

Adding retry logic in the Data Factory pipeline for the stored procedure activity is the least invasive fix. Deadlocks are transient by nature — SQL Server chooses a victim and rolls back its transaction — so retrying the activity typically succeeds on the next attempt without changing isolation levels or query semantics.

Exam trap

DP-203 often tests the misconception that disabling RCSI or adding NOLOCK will fix deadlocks — in reality deadlocks are transient and the correct minimal-impact fix is retry logic, not isolation-level changes.

How to eliminate wrong answers

Option A is wrong because NOLOCK hints introduce dirty reads and do not prevent deadlocks; they only reduce shared locks and can cause incorrect reporting data. Option B is wrong because removing the transaction breaks atomicity across the multiple table updates, risking inconsistent reporting data. Option D is wrong because disabling Read Committed Snapshot Isolation would revert to locking reads, increasing blocking and deadlock likelihood rather than reducing it.

63
MCQhard

You are implementing dynamic data masking on an Azure Synapse Analytics dedicated SQL pool. A table named Customers contains columns: CustomerID (int), Email (varchar), Phone (varchar), and CreditCard (varchar). You need to mask the Email and Phone columns so that users without elevated permissions see only the last four characters of the Email and Phone, while users with elevated permissions see the full values. You also need to ensure that the masking does not affect the storage size of the columns. What should you do?

A.Implement row-level security (RLS) on the Customers table to filter rows based on user permissions.
B.Encrypt the Email and Phone columns using Always Encrypted with deterministic encryption and grant decryption keys only to elevated users.
C.Create a masked view that concatenates '****' with the last four characters of Email and Phone, and grant SELECT on the view to users.
D.Use the MASKED WITH (FUNCTION = 'partial(0,"****",4)') clause on the Email and Phone columns and grant UNMASK permission to elevated users.
AnswerD

Dynamic data masking with the partial function allows you to mask a portion of the data while revealing the last four characters. The MASKED WITH clause is applied at the column level without changing storage size. Granting UNMASK permission to elevated users allows them to see the full data. This meets all requirements: masking for regular users, full access for privileged users, and no storage impact.

Why this answer

Dynamic data masking in Azure Synapse Analytics dedicated SQL pools allows you to define masking rules at the column level using the MASKED WITH clause. The partial function can reveal the last four characters while masking the rest. This does not change the underlying storage size because masking is applied at query time.

Granting UNMASK permission to privileged users allows them to see the full data, fulfilling the access requirements.

Exam trap

The trap here is confusing dynamic data masking with encryption or row-level security, which serve different purposes and do not provide partial masking without storage impact.

64
Multi-Selecteasy

Which TWO methods can you use to optimize the cost of storing data in Azure Data Lake Storage Gen2?

Select 2 answers
A.Use customer-managed keys for encryption.
B.Configure lifecycle management policies to move older data to the cool or archive tier.
C.Enable soft delete for blobs.
D.Use Azure Blob Storage access tiers: hot, cool, and archive.
E.Enable geo-redundant storage (GRS) for disaster recovery.
AnswersB, D

Lifecycle management policies automatically transition blobs between hot, cool and archive tiers based on age or last-access rules, directly reducing storage cost for ageing data. This satisfies the requirement to optimise Data Lake Storage Gen2 costs without manual intervention.

Why this answer

Option B is correct because Azure Data Lake Storage Gen2 lifecycle management policies automatically transition blobs to cooler access tiers (cool or archive) or delete them based on age or last-access rules, directly lowering storage costs for infrequently accessed data. Option D is correct because ADLS Gen2 is built on Azure Blob Storage, so data can be stored in the hot, cool, or archive access tiers, and choosing the appropriate tier for each dataset's access pattern minimizes per-GB storage charges. Option A is incorrect because customer-managed keys affect encryption key management and security/compliance, not storage cost.

Option C is incorrect because soft delete retains deleted blobs for a retention period, which adds storage consumption rather than reducing cost. Option E is incorrect because geo-redundant storage increases durability and disaster-recovery capability but raises cost compared with locally redundant storage.

65
MCQmedium

A company uses Azure Databricks for data processing. They want to monitor the performance of Spark jobs and set up alerts for job failures. Which Azure service should they use?

A.Azure Advisor
B.Azure Sentinel
C.Azure Log Analytics
D.Azure Monitor
AnswerD

Azure Monitor ingests Azure Databricks diagnostic logs and Spark metrics, enabling alert rules on job failures and performance thresholds. It is the native monitoring plane for the workspace, satisfying the requirement to track Spark job performance and trigger failure alerts.

Why this answer

Azure Monitor is the central service for collecting metrics and logs from Azure Databricks, enabling performance monitoring and alerting on Spark jobs. Option A is incorrect because Azure Advisor provides recommendations but not real-time monitoring or alerts. Option B is incorrect because Azure Sentinel is a SIEM solution for security incidents, not for job performance monitoring.

Option C is incorrect because Azure Log Analytics is a component of Azure Monitor used for log analysis, but the overarching service for monitoring and alerts is Azure Monitor.

Exam trap

Candidates often confuse Azure Monitor with Azure Log Analytics, but Log Analytics is a subset of Monitor. The question asks for the service to use, which is Azure Monitor.

66
MCQeasy

You need to monitor the health of your Azure Data Lake Storage Gen2 account. Which metric should you use to track the number of successful and failed requests?

A.Transactions.
B.Success E2E Latency.
C.Blob Capacity.
D.Ingress.
AnswerA

Transactions counts the total number of successful and failed requests against the storage account, directly satisfying the health-monitoring requirement. It is the only metric that surfaces both outcomes, letting you alert on failure spikes without enabling diagnostic logging.

Why this answer

Transactions metric tracks all requests. Option B is wrong because Success E2E Latency measures latency, not count. Option C is wrong because Blob Capacity measures storage size.

Option D is wrong because Ingress is about data incoming, not request count.

67
MCQmedium

You manage an Azure Data Lake Storage Gen2 account containing a large volume of JSON logs. Users frequently query only the last seven days of data, but each query scans the entire dataset, causing high costs and slow response times. You need to reduce the amount of data scanned by queries without changing the data format or moving the data. What should you do?

A.Partition the data by date into separate folders and update queries to filter on the date path.
B.Enable hierarchical namespace on the storage account.
C.Convert the JSON files to Parquet format.
D.Increase the throughput of the storage account.
AnswerA

Partitioning data by date into folders, such as year=2024/month=10/day=01, enables query engines to skip irrelevant folders when a query includes a filter on the partition column. This reduces the amount of data scanned, lowering cost and improving performance. Because the data format remains JSON and the files stay in the same account, no format conversion or data movement is required.

Why this answer

Partitioning the dataset by date and aligning queries with the partition path allows the query engine to prune unnecessary folders, dramatically reducing the bytes scanned. This approach keeps the original JSON format and location, satisfying the constraints. It directly addresses the root cause: lack of data organization that forces full scans.

Exam trap

The trap here is assuming that enabling hierarchical namespace or converting to a columnar format automatically reduces scanned data, when the real benefit comes from physically partitioning the data and filtering on the partition key.

68
MCQeasy

Your organization needs to ensure that all data stored in Azure Data Lake Storage Gen2 is encrypted at rest using Microsoft-managed keys. What is the default encryption method?

A.Storage Service Encryption (SSE) with Microsoft-managed keys.
B.Transparent Data Encryption (TDE) on the storage account.
C.Client-side encryption with keys stored in Azure Key Vault.
D.Azure Disk Encryption on the storage nodes.
AnswerA

Storage Service Encryption is enabled by default on every Azure Data Lake Storage Gen2 account, encrypting data at rest with Microsoft-managed keys automatically. No configuration is required, satisfying the requirement for Microsoft-managed key encryption without customer key setup.

Why this answer

Azure Data Lake Storage Gen2 is built on Azure Blob Storage, which automatically encrypts all data at rest using Storage Service Encryption (SSE) with Microsoft-managed keys by default. This encryption is always enabled and cannot be disabled. Microsoft-managed keys are used unless you choose to use customer-managed keys.

Exam trap

DP-203 often tests the default encryption method for storage services, and candidates may confuse TDE (for databases) or client-side encryption with the default SSE, leading to incorrect answers.

How to eliminate wrong answers

Option B is wrong because Transparent Data Encryption (TDE) is used for Azure SQL Database and SQL Managed Instance, not for storage accounts. Option C is wrong because client-side encryption is an optional feature where you encrypt data before uploading, not the default. Option D is wrong because Azure Disk Encryption is for VM disks, not for Data Lake Storage.

69
MCQmedium

Your organization uses Azure Synapse Analytics dedicated SQL pool. You need to ensure that all data at rest in the SQL pool is encrypted using a customer-managed key stored in Azure Key Vault. What should you configure?

A.Implement Always Encrypted with column encryption keys stored in Azure Key Vault.
B.Configure Dynamic Data Masking to obfuscate sensitive data.
C.Enable Azure Storage Service Encryption with a customer-managed key.
D.Enable Transparent Data Encryption (TDE) with a customer-managed key in Azure Key Vault.
AnswerD

Transparent Data Encryption operates at the storage layer, encrypting data and log files at rest, and supports a customer-managed key held in Azure Key Vault — satisfying the stem's requirement that the dedicated SQL pool's data at rest be encrypted under organisational key control.

Why this answer

Transparent Data Encryption (TDE) with a customer-managed key in Azure Key Vault encrypts the dedicated SQL pool's data at rest (data files, log files, backups) and allows the organization to control and rotate the encryption key. TDE is the native at-rest encryption mechanism for Azure Synapse dedicated SQL pools. Configuring it with a customer-managed key in Key Vault meets the requirement precisely.

Exam trap

DP-203 often tests the confusion between TDE (at-rest encryption of the whole database), Always Encrypted (column-level, client-side), and Storage Service Encryption (storage account level) — candidates must match the encryption scope to the requirement.

How to eliminate wrong answers

Option A is wrong because Always Encrypted protects individual columns from the database engine itself and is designed for client-side encryption of specific sensitive columns, not for encrypting all data at rest in the pool. Option B is wrong because Dynamic Data Masking only obfuscates data in query results for non-privileged users; it does not encrypt data at rest. Option C is wrong because Azure Storage Service Encryption applies to Azure Storage accounts (blobs, files), not to the dedicated SQL pool's internal storage, which is managed by the SQL engine.

70
MCQhard

You have an Azure Data Lake Storage Gen2 account that contains sensitive data. You need to implement a solution that enforces access control at the file and folder level, and also allows you to audit access. You want to minimize administrative effort. What should you do?

A.Use shared access signatures (SAS) with specific permissions and IP restrictions, and enable Azure Defender for Storage.
B.Configure Azure Private Endpoints and enable firewall rules to restrict access to the storage account.
C.Enable hierarchical namespace and configure POSIX access control lists (ACLs) on files and folders, and enable Azure Storage logging to capture access.
D.Enable Azure role-based access control (RBAC) on the storage account and assign the Storage Blob Data Contributor role to users.
AnswerC

Hierarchical namespace enables file and folder level ACLs, allowing granular permissions similar to a file system. Azure Storage logging (or Azure Monitor integration) can capture access details for auditing. This combination provides fine-grained access control and auditing with minimal administrative effort, as ACLs can be inherited and managed via tools like Azure Storage Explorer or scripts.

Why this answer

To enforce access control at the file and folder level, you need to enable hierarchical namespace and use POSIX ACLs. This allows granular permissions on directories and files. To audit access, enable Azure Storage logging or integrate with Azure Monitor.

This solution minimizes administrative effort because ACLs can be inherited and managed centrally. The other options either provide only coarse-grained access control or focus on network security rather than access auditing.

Exam trap

The trap here is assuming that Azure RBAC or SAS tokens provide file-level access control, when they are either too coarse or not designed for persistent granular access.

71
MCQmedium

You have an Azure Synapse Analytics dedicated SQL pool that handles both high-priority real-time queries and low-priority batch jobs. You need to ensure that high-priority queries always get the resources they need, while batch jobs do not starve. What should you configure?

A.Enable result-set caching for the high-priority queries
B.Enable data compression on the tables used by batch jobs
C.Create workload groups for high-priority and low-priority queries, assigning appropriate importance and resource percentages
D.Create materialized views for the batch job queries
AnswerC

Workload groups let you assign importance and a resource percentage per group, so high-priority queries pre-empt batch work while the batch group retains a guaranteed minimum, preventing starvation under the dedicated SQL pool's concurrency limits.

Why this answer

Workload groups in a dedicated SQL pool let you classify requests into groups and assign each group an importance level (e.g., HIGH for real-time queries, LOW for batch) plus a resource percentage (min/max memory and concurrency). This ensures high-priority queries preempt lower-priority ones for resources while guaranteeing batch jobs a floor so they don't starve. It is the native workload management mechanism for dedicated SQL pools.

Exam trap

DP-203 often tests the misconception that performance features like caching, compression, or materialized views solve resource contention, when only workload groups with importance actually govern prioritization.

How to eliminate wrong answers

Option A is wrong because result-set caching only speeds up repeated identical queries; it does not allocate or prioritize resources between competing workloads. Option B is wrong because data compression reduces storage and I/O but does not govern concurrency or memory allocation between high- and low-priority queries. Option D is wrong because materialized views precompute results for specific queries; they improve performance for those queries but do not implement resource governance or prevent batch jobs from starving real-time queries.

72
MCQmedium

You are a data engineer at a large retail company. Your team uses an Azure Synapse Analytics workspace with a dedicated SQL pool. You need to implement row-level security (RLS) so that sales representatives can see only data for their own region. You must ensure that the security predicate is evaluated at query time and that users cannot bypass it by using different tools. What should you do?

A.Create a database role, add users to it, and use a security policy with an inline table-valued function that filters rows based on the user's region.
B.Create a view that filters rows based on the CURRENT_USER function, and grant users SELECT permission only on the view.
C.Use dynamic data masking on the region column with a masking rule that shows only the user's region and masks others.
D.Implement column-level encryption on the region column and provide each sales representative with the encryption key for their region.
AnswerA

This is the correct approach for implementing row-level security in Azure Synapse Analytics dedicated SQL pool. You create an inline table-valued function that returns 1 when a user's region matches the row's region, then bind it to a security policy on the target table. Users are assigned to database roles, and the predicate is applied automatically to all queries, regardless of the client tool, ensuring consistent enforcement.

Why this answer

Row-level security in Azure Synapse Analytics dedicated SQL pool is implemented by creating an inline table-valued function that defines the filter predicate, then creating a security policy that binds that function to the table. Users are granted access via database roles. The predicate is applied transparently to all queries, ensuring that sales representatives can only see rows for their region, regardless of the client tool used.

This provides centralized, consistent enforcement.

Exam trap

The trap here is confusing row-level security with dynamic data masking or views, which do not filter rows at the engine level and can be bypassed or do not restrict row visibility.

73
MCQmedium

Your team is using Azure Synapse Analytics to process sensitive customer data. You need to ensure that column-level security is applied to a specific table so that only users with a certain role can view certain columns. Which feature should you use?

A.Column-level security (CLS)
B.Row-level security (RLS)
C.Azure Purview data policies
D.Dynamic data masking (DDM)
AnswerA

Column-level security uses GRANT and DENY statements on individual columns, restricting which roles can read specified fields within the table. This directly enforces the requirement that only users holding a certain role view particular columns.

Why this answer

Column-level security (CLS) in Azure Synapse Analytics allows you to restrict access to specific columns in a table based on the user's role or permissions. It is implemented using GRANT and DENY statements on individual columns, ensuring that only authorized users can view sensitive columns. This directly addresses the requirement to apply column-level security to a specific table.

Exam trap

DP-203 often tests the confusion between column-level security and dynamic data masking, where candidates might think DDM restricts access, but it only masks data and does not prevent retrieval of the actual values.

How to eliminate wrong answers

Option B is wrong because row-level security (RLS) filters rows based on user context, not columns, and is used to restrict access to specific rows rather than columns. Option C is wrong because Azure Purview data policies are used for data governance, classification, and access policies across the data estate, but they do not enforce column-level security within a Synapse table. Option D is wrong because dynamic data masking (DDM) obfuscates data in query results but does not prevent users from accessing the actual data; it is a masking technique, not an access control mechanism.

74
Multi-Selectmedium

You are optimizing an Azure Synapse Analytics dedicated SQL pool that stores a large fact table. Queries frequently join the fact table to a small dimension table on a non-distributed column, causing data movement. You need to reduce data movement and improve query performance. (Choose two.)

Select 2 answers
A.Replicate the small dimension table so a full copy exists on every compute node.
B.Change the distribution of the large fact table to ROUND_ROBIN.
C.Hash-distribute the large fact table on the column used in the join predicate.
D.Partition the large fact table by date and rebuild statistics.
E.Create a clustered columnstore index on the small dimension table.
AnswersA, C

Replicating the small dimension table places a full copy on each compute node. When the fact table is hash-distributed on the join key, the join can occur locally on each node without shuffling the dimension data. This eliminates data movement for the dimension side, which is a standard optimization for star-schema joins in dedicated SQL pools.

Why this answer

To eliminate data movement for a join between a large fact table and a small dimension, align the fact table's distribution with the join key and replicate the dimension. Hash-distributing the fact table on the join column co-locates matching rows, while replication places the small dimension on every node. Together they enable local joins without shuffling large datasets across the fabric.

Exam trap

The trap here is assuming that partitioning, statistics, or columnstore indexes eliminate data movement, when only distribution alignment and replication address the shuffle caused by joins on non-distributed columns.

75
MCQmedium

You are a data engineer at a logistics company. You have an Azure Data Lake Storage Gen2 account that stores JSON logs from IoT devices. The logs are written continuously and are stored in a folder structure of /logs/{year}/{month}/{day}/{hour}/. You need to optimize the storage for cost and performance. The data is accessed frequently for the first 30 days, then occasionally for the next 60 days, and rarely after that. You need to minimize storage costs while ensuring that data remains available. What should you do?

A.Enable soft delete for blobs and set the retention period to 90 days.
B.Create a scheduled Azure Data Factory pipeline that moves data older than 30 days to a separate storage account with the Cool tier, and older than 90 days to the Archive tier.
C.Configure a lifecycle management policy to move blobs to the Cool tier after 30 days and to the Archive tier after 90 days.
D.Use Azure Data Lake Storage Gen2 hierarchical namespace and set POSIX permissions to restrict access to older data.
AnswerC

Azure Blob Storage lifecycle management policies can automatically transition blobs between access tiers based on age. Moving data to Cool after 30 days reduces storage costs for infrequently accessed data, and moving to Archive after 90 days further reduces costs for rarely accessed data. This aligns with the access pattern and minimizes costs while keeping data available (though Archive requires rehydration for access).

Why this answer

Lifecycle management policies in Azure Storage automatically transition blobs between Hot, Cool, and Archive tiers based on rules you define. This matches the access pattern: frequent access for 30 days (Hot), occasional for next 60 days (Cool), and rare thereafter (Archive). It minimizes storage costs without manual intervention, and data remains available (Archive requires rehydration).

This is the most efficient and cost-effective solution.

Exam trap

The trap here is confusing data protection features like soft delete or access control mechanisms with cost optimization, when the requirement is specifically about automatically moving data to cheaper tiers based on age.

Page 1 of 3 · 159 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Secure, monitor, and optimize data storage and data processing questions.