Courseiva

DP-203 Design and implement data storage Practice Question

You are designing a storage layer in Azure Data Lake Storage Gen2 for a data engineering pipeline. You need to store Parquet files that will be queried by Azure Synapse Analytics serverless SQL pools and Azure Databricks. You must optimize for query performance and minimize data scanned. Which two actions should you perform? (Choose two.)

⚠ Common exam trap

The trap here is assuming that any partitioning improves performance, when high-cardinality partitioning creates many small files and degrades it.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Partition the data by a commonly filtered column such as TransactionDate.

Using Parquet with a columnar layout enables column pruning and compression, reducing the data scanned. Partitioning by a commonly filtered column such as TransactionDate allows partition elimination, further reducing the amount of data read. Together, these actions optimize query performance for both Synapse serverless SQL pools and Databricks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Partition the data by a commonly filtered column such as TransactionDate.

    Why this is correct

    Partitioning by a frequently filtered column allows queries with predicates on that column to skip entire partitions, reducing the amount of data scanned. This is especially effective when the column has a moderate number of distinct values, such as dates. It aligns with the goal of minimizing data scanned and improving performance in both Synapse and Databricks.

  • ✗

    Store the data in CSV format to ensure compatibility with all query engines.

    Why it's wrong here

    CSV is a row-based text format that lacks column pruning and efficient compression, so queries scan all columns and more data than necessary. While CSV is widely compatible, it does not optimize for performance or minimize data scanned. Using CSV would contradict the requirement to optimize for query performance in analytical engines.

  • ✗

    Partition the data by a high-cardinality column such as TransactionId.

    Why it's wrong here

    Partitioning by a high-cardinality column like TransactionId creates a very large number of small files and directories, which increases metadata overhead and reduces query performance. Partitioning should be done on columns with moderate cardinality that are commonly used in filters, such as date or region. This action would degrade rather than improve performance.

  • ✗

    Enable hierarchical namespace and store all files in a single directory.

    Why it's wrong here

    Hierarchical namespace is a feature of ADLS Gen2 that enables directory semantics, but storing all files in a single directory does not provide partitioning benefits. Without a folder structure that reflects common filters, queries cannot prune data. This action does not minimize data scanned and may actually hurt performance due to large directory listings.

  • ✓

    Store the data in Parquet format with a columnar layout.

    Why this is correct

    Parquet is a columnar format that enables column pruning and efficient compression, so queries that select a subset of columns scan less data. Both Synapse serverless SQL pools and Databricks can read Parquet efficiently and push down filters. This directly minimizes the amount of data scanned and improves query performance for analytical workloads.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

Go deeper

Related to this question

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.