Courseiva

DP-203 · topic practice

Design and implement data storage practice questions

This domain covers choosing and configuring Azure storage for analytical workloads: Azure Data Lake Storage Gen2, Blob Storage tiers, Azure Synapse Analytics pools, and Azure Cosmos DB. Questions present a scenario with access patterns, latency, format, and cost constraints, then ask you to select the service, storage layer, or partitioning strategy that satisfies all stated requirements.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Design and implement data storage

What the exam tests

What to know about Design and implement data storage

Map each scenario's access frequency, latency, format, and cost constraints to the correct Azure storage service, tier, and partitioning scheme. The single most important thing: verify every stated requirement is satisfied, since one unmet constraint (like latency or immutability) eliminates an otherwise attractive option.

Selecting ADLS Gen2 hierarchical namespace versus flat Blob Storage for analytical data lakes

Choosing Blob access tiers (hot, cool, archive) and lifecycle management rules for cost

Designing Synapse dedicated SQL pool distribution (hash, round-robin, replicated) and partitioning

Configuring Cosmos DB partitioning, consistency levels, and analytical store for mixed workloads

Watch out for

Common Design and implement data storage exam traps

  • ▸Assuming archive tier supports immediate reads; rehydration takes hours and is billed, so it fails low-latency requirements.
  • ▸Ignoring that ADLS Gen2 hierarchical namespace is required for Synapse and Databricks directory-level operations and POSIX ACLs.
  • ▸Picking round-robin distribution for large fact tables that are frequently joined, causing costly data movement instead of hash distribution.

Practice set

Design and implement data storage questions

20 questions · select your answer, then reveal the explanation

You are designing a solution to store streaming data from multiple sources into Azure Data Lake Storage Gen2. The data must be organized by ingestion time and source system. Each source system produces data in a different format: CSV, JSON, and Parquet. The solution must allow efficient querying using Azure Synapse Serverless SQL and must support partitioning on ingestion date. What is the recommended folder structure?

A healthcare company stores sensitive patient data in Azure Data Lake Storage Gen2. They need to ensure that only authorized users can access data and that all access is audited. They also need to prevent data from being accessed by unauthorized Azure services. Which combination of security features should be used?

Which THREE of the following are required to configure a managed private endpoint for Azure Data Factory when connecting to an Azure SQL Database that has a private endpoint?

You are reviewing a copy job configuration in Azure Data Factory that copies Parquet files from Azure Data Lake Storage Gen2 to Azure Synapse Analytics. The exhibit shows the job settings. If the source folder contains a file that is not in Parquet format (e.g., a CSV file), what will happen?

Exhibit

Refer to the exhibit.

```json
{
  "data": [
    {
      "name": "order_data",
      "path": "orders/*.parquet",
      "partitionBy": ["year", "month", "day"],
      "format": "parquet",
      "options": {
        "compression": "snappy"
      }
    }
  ],
  "source": {
    "provider": "AzureDataLakeStorage",
    "connectionString": "DefaultEndpointsProtocol=https;AccountName=storagedatalake;AccountKey=...;EndpointSuffix=core.windows.net",
    "container": "data"
  },
  "sink": {
    "provider": "AzureSynapseAnalytics",
    "table": "dbo.orders",
    "staging": {
      "linkedServiceName": "AzureDataLakeStorage",
      "folderPath": "staging"
    }
  },
  "copyBehavior": "MergeFiles",
  "faultTolerance": {
    "skipIncompatibleFiles": true,
    "skipIncompatibleRows": true
  }
}
```

A company stores sensitive customer data in Azure Data Lake Storage Gen2. They need to implement a data retention policy where data older than 90 days is automatically moved to the 'cold' access tier, and data older than 365 days is deleted. Which Azure feature should be used to automate this?

A company ingests streaming data from multiple sources into Azure Event Hubs. The data must be stored in Azure Data Lake Storage Gen2 in Parquet format, partitioned by date and hour. The solution must minimize cost and processing latency. Which THREE actions should be taken?

A company stores IoT sensor data in Azure Blob Storage. The data is appended every minute and must be queried in near real-time using a SQL interface. Which Azure service should be used to enable this?

A company is designing a data lake on Azure Data Lake Storage Gen2. Data comes from multiple sources with varying schemas. The team must minimize storage costs while keeping all data available for future processing. Which storage tier should they use for the raw ingested data?

You are designing a solution to store telemetry data from millions of devices. Each device sends a JSON payload every 5 seconds. The data must be partitioned by device ID and time for efficient querying and must support real-time streaming ingestion. Which Azure storage solution should you recommend?

A company uses Azure SQL Database for an OLTP application. They need to run complex analytical queries without impacting OLTP performance. Which solution should they implement?

Which THREE security features are available for Azure SQL Database? (Choose three.)

A data engineer runs the Azure CLI command shown in the exhibit. The blob is stored in Azure Blob Storage. The team previously set a lifecycle management rule to move blobs to the Archive tier after 30 days. The blob was created 45 days ago. What is the most likely reason the blob is still in the Cool tier?

Exhibit

Refer to the exhibit.

az storage blob show \
  --account-name exampledatalake \
  --container-name raw \
  --name sensor/2023/01/01/data.parquet \
  --query "properties.blobTier"

Output: "Cool"

You are a data engineer for a financial services company. The company stores sensitive transaction data in Azure Data Lake Storage Gen2. The data is partitioned by date and loaded daily via Azure Data Factory. Recently, an audit found that the storage account allows public network access, and some containers have anonymous read access enabled. You need to secure the storage account according to the principle of least privilege while ensuring that Azure Data Factory can still load data. You must also ensure that data can be accessed by Azure Databricks for analytics. The solution must minimize administrative overhead. Which course of action should you take?

A retail company uses Azure Synapse Analytics dedicated SQL pool to store sales data. The data is loaded nightly from Azure Data Lake Storage Gen2 using PolyBase. Recently, the load process started failing with the error 'External table 'sales' is not accessible because the location does not exist or is used by another process.' You verify that the storage account, container, and file path are correct. The file is a CSV file named 'sales_20250301.csv' and it exists. Other files in the same container load successfully. What is the most likely cause of the error?

A healthcare company stores patient records in Azure Blob Storage. The compliance team requires that all data be encrypted at rest using customer-managed keys (CMK) stored in Azure Key Vault. Additionally, the storage account must be accessible only from a specific virtual network (VNet) and must support versioning to protect against accidental deletion. The storage account is currently using Microsoft-managed keys and has public network access enabled. You need to implement the required changes with minimal downtime. Which course of action should you take?

A company is migrating an on-premises Hadoop cluster to Azure. The cluster uses Hive tables stored as Parquet files on HDFS. They want to minimize changes to existing Hive queries and continue using HiveQL. Which Azure storage solution should they choose?

Which TWO of the following are recommended practices for designing a data storage solution using Azure Data Lake Storage Gen2?

Which THREE of the following are valid methods to load data into Azure Synapse Analytics?

You need to partition a large Azure SQL Database table by date to improve query performance and manageability. Which partitioning strategy should you use?

Match each Azure data storage service to its primary use case.

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Design and implement data storage sessions

Start a Design and implement data storage only practice session

Every question in these sessions is drawn from the Design and implement data storage domain — nothing else.

Related practice questions

Related DP-203 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the DP-203 exam test about Design and implement data storage?
Map each scenario's access frequency, latency, format, and cost constraints to the correct Azure storage service, tier, and partitioning scheme. The single most important thing: verify every stated requirement is satisfied, since one unmet constraint (like latency or immutability) eliminates an otherwise attractive option.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Design and implement data storage questions in a focused session?
Yes — the session launcher on this page draws every question from the Design and implement data storage domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other DP-203 topics?
Use the topic links above to move to related areas, or go back to the DP-203 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the DP-203 exam covers. They are not copied from any real exam or dump site.