Courseiva
← Back to Microsoft Azure Data Engineer Associate DP-203 questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise Microsoft Azure Data Engineer Associate DP-203 practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

15
scenario questions
DP-203
exam code
Microsoft
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related DP-203 topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

Refer to the exhibit. You are creating a serverless SQL table in Azure Synapse Analytics that reads Parquet files from the specified location. The folder contains multiple Parquet files with different schemas. When querying the table, you get an error about schema mismatch. What is the most likely reason?

Exhibit

{
  "type": "Microsoft.Synapse/workspaces/databases/tables",
  "properties": {
    "source": {
      "provider": "ABFS",
      "location": "abfss://container@storage.dfs.core.windows.net/data/"
    },
    "format": {
      "type": "parquet",
      "derivedModel": false
    },
    "options": {
      "recursive": true
    }
  }
}
Question 2easymultiple choice
Full question →

Refer to the exhibit. An Azure Policy is defined to enforce network security on storage accounts. What does this policy do?

Exhibit

Refer to the exhibit.

{
  "policyRule": {
    "if": {
      "field": "type",
      "equals": "Microsoft.Storage/storageAccounts"
    },
    "then": {
      "effect": "deny",
      "details": {
        "field": "Microsoft.Storage/storageAccounts/networkAcls.defaultAction",
        "equals": "Allow"
      }
    }
  }
}
Question 3hardmultiple choice
Full question →

Refer to the exhibit. You are reviewing an ARM template for an Azure Data Lake Storage Gen2 account. Which of the following security best practices is violated in this template?

Exhibit

Refer to the exhibit.

{
  "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#",
  "contentVersion": "1.0.0.0",
  "resources": [
    {
      "type": "Microsoft.Storage/storageAccounts",
      "apiVersion": "2022-09-01",
      "name": "[parameters('storageAccountName')]",
      "location": "[resourceGroup().location]",
      "kind": "StorageV2",
      "sku": {
        "name": "Standard_LRS"
      },
      "properties": {
        "supportsHttpsTrafficOnly": false,
        "minimumTlsVersion": "TLS1_0",
        "isHnsEnabled": true
      }
    }
  ]
}
Question 4easymultiple choice
Full question →

Refer to the exhibit. You are reviewing an Azure Stream Analytics job query. The job has a stream input and a reference data input. The job is failing with the error 'Reference data input must be of type Reference, not Stream'. What is the cause of the error?

Exhibit

Refer to the exhibit.

{
  "inputAlias": "input",
  "type": "Stream",
  "output": {
    "outputAlias": "output",
    "type": "ReferenceData"
  },
  "query": "SELECT input.* FROM input JOIN output ON input.ProductId = output.ProductId"
}
Question 5hardmultiple choice
Full question →

Refer to the exhibit. A data engineer notices that the copy activity sometimes copies 0 rows despite reading 1 million rows. What is the most likely cause?

Network Topology
|Metrics from Azure Monitor:
Question 6mediummultiple choice
Full question →

Refer to the exhibit. You are deploying an Azure Synapse Analytics workspace using an ARM template. The exhibit shows the encryption configuration. What is the effect of setting infrastructureEncryption to Enabled?

Exhibit

Refer to the exhibit.

{
  "identity": {
    "type": "SystemAssigned",
    "principalId": "00000000-0000-0000-0000-000000000001",
    "tenantId": "00000000-0000-0000-0000-000000000002"
  },
  "properties": {
    "encryption": {
      "keyVaultProperties": {
        "keyVaultUri": "https://mykeyvault.vault.azure.net/",
        "keyName": "mykey",
        "keyVersion": "1"
      },
      "infrastructureEncryption": "Enabled"
    }
  }
}
Question 7mediummultiple choice
Full question →

Refer to the exhibit. An ARM template deploys an Azure Synapse Analytics workspace. What is the purpose of the 'managedVirtualNetwork' property set to 'default'?

Exhibit

Refer to the exhibit.

{
  "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#",
  "contentVersion": "1.0.0.0",
  "resources": [
    {
      "type": "Microsoft.Synapse/workspaces",
      "apiVersion": "2021-06-01",
      "name": "myworkspace",
      "location": "eastus",
      "properties": {
        "defaultDataLakeStorage": {
          "accountUrl": "https://mystorageaccount.dfs.core.windows.net",
          "filesystem": "synapse"
        },
        "sqlAdministratorLogin": "admin",
        "sqlAdministratorLoginPassword": "P@ssw0rd123!",
        "managedVirtualNetwork": "default"
      },
      "identity": {
        "type": "SystemAssigned"
      }
    }
  ]
}
Question 8hardmultiple choice
Full question →

Refer to the exhibit. You are deploying an Azure Synapse Analytics dedicated SQL pool using the provided ARM template snippet. After deployment, you need to adjust the performance level to DW200c to handle increased workload. Which parameter should you modify?

Exhibit

Refer to the exhibit.

{
  "type": "Microsoft.Synapse/workspaces/sqlPools",
  "properties": {
    "createMode": "Default",
    "storageAccountType": "GRS",
    "collation": "SQL_Latin1_General_CP1_CI_AS",
    "maxSizeBytes": 263882790666240,
    "sku": {
      "name": "DW100c",
      "tier": "ServiceObjective"
    },
    "restorePointInTime": null,
    "sourceDatabaseId": null
  }
}
Question 9hardmultiple choice
Full question →

Refer to the exhibit. A Bicep file is used to deploy an Azure Synapse Analytics workspace. What is the purpose of the 'purviewConfiguration' property?

Exhibit

Refer to the exhibit.

{
  "properties": {
    "dataLakeStorageAccountDetails": [
      {
        "accountUrl": "https://mystorageaccount.dfs.core.windows.net"
      }
    ],
    "defaultDataLakeStorage": {
      "accountUrl": "https://mystorageaccount.dfs.core.windows.net",
      "filesystem": "synapseworkspace"
    },
    "sqlAdministratorLogin": "adminuser",
    "sqlAdministratorLoginPassword": "",
    "managedResourceGroupName": "managedRG",
    "purviewConfiguration": {
      "purviewResourceId": "/subscriptions/sub-id/resourceGroups/rg/providers/Microsoft.Purview/accounts/purview-account"
    },
    "encryption": {
      "cmk": {
        "key": {
          "name": "cmk-key",
          "keyVaultUrl": "https://kv.vault.azure.net/"
        }
      }
    }
  }
}
Question 10hardmultiple choice
Full question →

Refer to the exhibit. You submit a Spark job in Azure Synapse Analytics using the Azure CLI. The job runs slowly during the shuffle phase. The input data is about 200 GB. Which configuration change would best improve performance for this shuffle-heavy workload?

Network Topology
name MyJobfile abfss://container@storage.dfs.core.windows.net/path/etl.pyexecutor-size Smallexecutors 2conf spark.sql.shuffle.partitions=400"
Question 11hardmultiple choice
Full question →

Refer to the exhibit. A Stream Analytics job shows increasing watermark delay and input deserialization errors. Which action should be taken first to troubleshoot?

Exhibit

Azure Stream Analytics job diagnostics log:

{
  "time": "2023-08-01T12:00:00Z",
  "properties": {
    "jobId": "job-123",
    "jobName": "IoTStreamJob",
    "events": [
      {
        "time": "2023-08-01T11:59:00Z",
        "type": "WatermarkDelay",
        "properties": {
          "watermarkDelaySeconds": 120,
          "maxWatermarkDelaySeconds": 300
        }
      },
      {
        "time": "2023-08-01T11:59:30Z",
        "type": "InputDeserializationError",
        "properties": {
          "source": "iothub",
          "count": 15
        }
      }
    ],
    "jobOutputWatermark": "2023-08-01T11:57:00Z"
  }
}
Question 12mediummultiple choice
Full question →

Refer to the exhibit. You are reviewing a Data Factory JSON definition. The factory has a user-assigned managed identity configured. However, the linked service to Azure Storage uses an account key. What security improvement should you recommend?

Exhibit

Refer to the exhibit.

{
  "identity": {
    "type": "UserAssigned",
    "userAssignedIdentities": {
      "/subscriptions/.../resourceGroups/rg1/providers/Microsoft.ManagedIdentity/userAssignedIdentities/mi-etl": {}
    }
  },
  "properties": {
    "linkedServices": [
      {
        "name": "ls_storage",
        "type": "AzureStorage",
        "typeProperties": {
          "connectionString": "DefaultEndpointsProtocol=https;AccountName=mystorage;AccountKey=mykey"
        }
      }
    ]
  }
}
Question 13mediummultiple choice
Full question →

Refer to the exhibit. A data engineer wants to copy only new orders from an Azure SQL database to Azure Data Lake Storage Gen2. The pipeline runs daily at midnight. What should be added to the pipeline to ensure incremental loads?

Exhibit

{
  "name": "CopyDataPipeline",
  "properties": {
    "activities": [
      {
        "name": "CopyData",
        "type": "Copy",
        "inputs": [{"name": "InputDataset"}],
        "outputs": [{"name": "OutputDataset"}],
        "typeProperties": {
          "source": {
            "type": "AzureSqlSource",
            "sqlReaderQuery": "SELECT * FROM Orders WHERE OrderDate > '2024-01-01'"
          },
          "sink": {
            "type": "DelimitedTextSink",
            "storeSettings": {
              "type": "AzureBlobFSWriteSettings",
              "copyBehavior": "PreserveHierarchy"
            }
          },
          "enableStaging": false,
          "translator": {
            "type": "TabularTranslator",
            "columnMappings": {
              "OrderID": "order_id",
              "CustomerID": "customer_id",
              "OrderDate": "order_date"
            }
          }
        }
      }
    ]
  }
}
Question 14easymultiple choice
Full question →

Refer to the exhibit. You have a mapping data flow in Azure Data Factory that aggregates sales data. The data flow runs successfully but the sink table contains only the total sum per run instead of per product. What is missing?

Exhibit

{
  "dataflows": [
    {
      "name": "TransformSales",
      "properties": {
        "sources": [
          {
            "name": "SalesSource",
            "dataset": {
              "referenceName": "SalesDataset",
              "type": "DatasetReference"
            }
          }
        ],
        "transformations": [
          {
            "name": "AggregateSales",
            "type": "Aggregate",
            "inputs": ["SalesSource"],
            "aggregates": [
              {
                "column": "TotalAmount",
                "function": "SUM",
                "input": "Amount"
              }
            ]
          }
        ],
        "sink": {
          "name": "SalesSink",
          "dataset": {
            "referenceName": "AggregatedSalesDataset",
            "type": "DatasetReference"
          }
        }
      }
    }
  ]
}
Question 15mediummultiple choice
Full question →

Refer to the exhibit. You are deploying an Azure Synapse Analytics workspace using an ARM template. The template defines a managed virtual network integration runtime. You need to ensure that the integration runtime can run mapping data flows with a time-to-live (TTL) of 10 minutes. What is the purpose of the 'timeToLive' property in this configuration?

Exhibit

Refer to the exhibit.

{
    "$schema": "https://schema.management.azure.com/schemas/2019-04-01/deploymentTemplate.json#",
    "contentVersion": "1.0.0.0",
    "resources": [
        {
            "type": "Microsoft.Synapse/workspaces/integrationRuntimes",
            "apiVersion": "2021-06-01-preview",
            "name": "[concat(parameters('workspaceName'), '/MyManagedVNetIR')]",
            "properties": {
                "type": "Managed",
                "typeProperties": {
                    "computeProperties": {
                        "location": "AutoResolve",
                        "dataFlowProperties": {
                            "computeType": "General",
                            "coreCount": 8,
                            "timeToLive": 10
                        }
                    }
                }
            }
        }
    ]
}

These DP-203 practice questions are part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style DP-203 questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.