DP-203 Design and implement data storage Practice Question
You are designing a storage layer in Azure Data Lake Storage Gen2 for a data engineering pipeline. You need to store Parquet files that will be queried by Azure Synapse Analytics serverless SQL pools and Azure Databricks. You must optimize for query performance and minimize data scanned. Which two actions should you perform? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any partitioning improves performance, when high-cardinality partitioning creates many small files and degrades it.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partition the data by a commonly filtered column such as TransactionDate.
Using Parquet with a columnar layout enables column pruning and compression, reducing the data scanned. Partitioning by a commonly filtered column such as TransactionDate allows partition elimination, further reducing the amount of data read. Together, these actions optimize query performance for both Synapse serverless SQL pools and Databricks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Partition the data by a commonly filtered column such as TransactionDate.
Why this is correct
Partitioning by a frequently filtered column allows queries with predicates on that column to skip entire partitions, reducing the amount of data scanned. This is especially effective when the column has a moderate number of distinct values, such as dates. It aligns with the goal of minimizing data scanned and improving performance in both Synapse and Databricks.
- ✗
Store the data in CSV format to ensure compatibility with all query engines.
Why it's wrong here
CSV is a row-based text format that lacks column pruning and efficient compression, so queries scan all columns and more data than necessary. While CSV is widely compatible, it does not optimize for performance or minimize data scanned. Using CSV would contradict the requirement to optimize for query performance in analytical engines.
- ✗
Partition the data by a high-cardinality column such as TransactionId.
Why it's wrong here
Partitioning by a high-cardinality column like TransactionId creates a very large number of small files and directories, which increases metadata overhead and reduces query performance. Partitioning should be done on columns with moderate cardinality that are commonly used in filters, such as date or region. This action would degrade rather than improve performance.
- ✗
Enable hierarchical namespace and store all files in a single directory.
Why it's wrong here
Hierarchical namespace is a feature of ADLS Gen2 that enables directory semantics, but storing all files in a single directory does not provide partitioning benefits. Without a folder structure that reflects common filters, queries cannot prune data. This action does not minimize data scanned and may actually hurt performance due to large directory listings.
- ✓
Store the data in Parquet format with a columnar layout.
Why this is correct
Parquet is a columnar format that enables column pruning and efficient compression, so queries that select a subset of columns scan less data. Both Synapse serverless SQL pools and Databricks can read Parquet efficiently and push down filters. This directly minimizes the amount of data scanned and improves query performance for analytical workloads.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
Learn chapter
Implement Azure Synapse Analytics
Key term
Azure Databricks
Azure Databricks is a fast, easy, and collaborative Apache Spark-based analytics platform optimized for Azure that lets data teams prepare data, run machine learning models, and build data pipelines using a single workspace.
Key term
Azure Synapse Analytics
Azure Synapse Analytics is a cloud-based data integration, warehousing, and analytics service that brings together big data and data warehouse capabilities under one platform.
About these practice questions
One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.