DP-203 Design and implement data storage Practice Question
You are designing a storage solution for a financial services company. The solution must store large volumes of semi-structured JSON data in Azure Data Lake Storage Gen2. The data is accessed by Azure Databricks for batch processing and by Azure Synapse Analytics for interactive queries. The data must be organized for efficient partition elimination and must support atomic operations. You need to choose the appropriate file format and partitioning strategy. What should you do?
⚠ Common exam trap
The trap here is assuming that any file format with partitioning will meet the performance requirements, but columnar formats like Parquet are essential for efficient analytical queries.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store the data as Parquet files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.
Parquet is the optimal format for analytical workloads due to its columnar storage, compression, and predicate pushdown capabilities. Partitioning by date enables partition elimination, which is critical for efficient querying. Registering the folder as an external table in both Azure Databricks and Azure Synapse Analytics allows both services to access the same data seamlessly. Other formats like CSV and Avro are not as efficient for these requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store the data as CSV files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.
Why it's wrong here
CSV is a row-based format and does not support efficient columnar compression or predicate pushdown as well as Parquet. While partitioning by date helps with partition elimination, the lack of columnar storage and schema evolution support makes CSV less efficient for analytical queries. Atomic operations are also more challenging with CSV because appending to a file is not atomic.
- ✗
Store the data as Parquet files without partitioning, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.
Why it's wrong here
Parquet is a good choice, but without partitioning, queries cannot take advantage of partition elimination. The requirement explicitly states that the data must be organized for efficient partition elimination. Without partitioning, all data would be scanned, leading to slower queries and higher costs. Partitioning by a common filter column like date is necessary to meet the requirement.
- ✗
Store the data as Avro files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.
Why it's wrong here
Avro is a row-based format and is more suitable for write-heavy workloads and schema evolution, but it is not optimized for analytical queries like Parquet. While partitioning by date helps with partition elimination, Avro does not provide the same columnar compression and predicate pushdown benefits. For interactive queries in Synapse, Parquet is the preferred format.
- ✓
Store the data as Parquet files partitioned by date, and register the folder as an external table in both Azure Databricks and Azure Synapse Analytics.
Why this is correct
Parquet is a columnar format ideal for analytical workloads, offering efficient compression and predicate pushdown. Partitioning by date enables partition elimination, reducing the amount of data scanned. Registering the folder as an external table in both services allows them to query the same data without duplication. This meets the requirements for efficient querying and atomic operations (Parquet files are immutable, and writes can be atomic at file level).
Go deeper
Related to this question
Learn chapter
Implement Azure Stream Analytics
Key term
Azure Databricks
Azure Databricks is a fast, easy, and collaborative Apache Spark-based analytics platform optimized for Azure that lets data teams prepare data, run machine learning models, and build data pipelines using a single workspace.
Key term
Azure Synapse Analytics
Azure Synapse Analytics is a cloud-based data integration, warehousing, and analytics service that brings together big data and data warehouse capabilities under one platform.
About these practice questions
One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.