Courseiva

DP-203 Design and implement data storage Practice Question

You are designing a storage solution for a financial services company. The solution must store large volumes of semi-structured JSON data in Azure Data Lake Storage Gen2. The data is accessed by Azure Databricks for batch processing and by Azure Synapse Analytics for interactive queries. The data must be organized for efficient partition elimination and must support atomic operations. You need to choose the appropriate file format and partitioning strategy. What should you do?

⚠ Common exam trap

The trap here is assuming that any file format with partitioning will meet the performance requirements, but columnar formats like Parquet are essential for efficient analytical queries.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Store the data as Parquet files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.

Parquet is the optimal format for analytical workloads due to its columnar storage, compression, and predicate pushdown capabilities. Partitioning by date enables partition elimination, which is critical for efficient querying. Registering the folder as an external table in both Azure Databricks and Azure Synapse Analytics allows both services to access the same data seamlessly. Other formats like CSV and Avro are not as efficient for these requirements.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Store the data as CSV files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.

    Why it's wrong here

    CSV is a row-based format and does not support efficient columnar compression or predicate pushdown as well as Parquet. While partitioning by date helps with partition elimination, the lack of columnar storage and schema evolution support makes CSV less efficient for analytical queries. Atomic operations are also more challenging with CSV because appending to a file is not atomic.

  • ✗

    Store the data as Parquet files without partitioning, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.

    Why it's wrong here

    Parquet is a good choice, but without partitioning, queries cannot take advantage of partition elimination. The requirement explicitly states that the data must be organized for efficient partition elimination. Without partitioning, all data would be scanned, leading to slower queries and higher costs. Partitioning by a common filter column like date is necessary to meet the requirement.

  • ✗

    Store the data as Avro files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.

    Why it's wrong here

    Avro is a row-based format and is more suitable for write-heavy workloads and schema evolution, but it is not optimized for analytical queries like Parquet. While partitioning by date helps with partition elimination, Avro does not provide the same columnar compression and predicate pushdown benefits. For interactive queries in Synapse, Parquet is the preferred format.

  • ✓

    Store the data as Parquet files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.

    Why this is correct

    Parquet is a columnar format ideal for analytical workloads, offering efficient compression and predicate pushdown. Partitioning by date enables partition elimination, reducing the amount of data scanned. Registering the folder as an external table in both services allows them to query the same data without duplication. This meets the requirements for efficient querying and atomic operations (Parquet files are immutable, and writes can be atomic at file level).

Go deeper

Related to this question

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.