DP-203 Design and implement data storage Practice Question
You are designing a storage layer for a fraud detection system. The system writes millions of small JSON records per hour to Azure Data Lake Storage Gen2 and must support both batch analytics and interactive queries from Azure Databricks. You need to choose a storage format and layout that minimizes query latency for selective filters on customer ID while keeping storage costs predictable. What should you do?
⚠ Common exam trap
The trap here is choosing Parquet with a high-cardinality partition column, which creates too many small partitions instead of using clustering within files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store the data as Delta Lake tables with Z-ORDER clustering on customer ID and optimize file sizes.
Delta Lake with Z-ORDER clustering on customer ID enables data skipping by storing min/max statistics per file and colocating similar values. This directly reduces the data read for selective filters. OPTIMIZE compacts small files, which is critical when millions of small records arrive hourly, improving both batch and interactive query performance while keeping storage costs controlled.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Store the data as Delta Lake tables with Z-ORDER clustering on customer ID and optimize file sizes.
Why this is correct
Delta Lake provides ACID transactions and metadata that Databricks can use for data skipping. Z-ORDER clustering colocates related customer ID values within files, so selective filters skip irrelevant data efficiently. Combined with OPTIMIZE to compact small files, this reduces latency and keeps storage costs predictable through compaction and retention policies, matching the scenario's needs.
- ✗
Store the data as Parquet files partitioned by customer ID and use partition pruning in Databricks.
Why it's wrong here
Partitioning by customer ID would create an extremely large number of small partitions, causing metadata overhead and slow listing operations. While Parquet is a good format, the partition key choice is poor for high-cardinality columns. This leads to the small-file problem, degrading performance instead of improving it for selective customer filters.
- ✗
Store the data as Avro files and use schema evolution to handle new fields.
Why it's wrong here
Avro is row-oriented and designed for write-heavy serialization, not for selective analytical reads. It does not support the file-level statistics and clustering that enable data skipping for customer ID filters. While schema evolution is useful, it does not address the latency or cost requirements, so this choice is unsuitable for interactive queries.
- ✗
Store the data as uncompressed JSON files and rely on Databricks caching to improve performance.
Why it's wrong here
Uncompressed JSON is row-oriented and expensive to scan for selective filters, and caching only helps after the first read. It does not reduce storage cost or improve initial query latency for customer-level lookups. This format also inflates storage footprint, making costs less predictable, so it fails the scenario's requirements for latency and predictable cost.
Go deeper
Related to this question
About these practice questions
One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.