Courseiva

DP-203 Design and implement data storage Practice Question

You are designing a storage solution for a financial analytics platform. The data consists of large Parquet files stored in Azure Data Lake Storage Gen2. Analysts run complex queries that scan entire partitions, but only a subset of columns is needed for each query. You need to minimize the amount of data read from storage and improve query performance. What should you do?

⚠ Common exam trap

The trap here is assuming that any compression or file organization automatically reduces data read during queries, when only columnar formats with column pruning achieve that.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Parquet format with column pruning and predicate pushdown enabled in the query engine.

Parquet is a columnar storage format that enables column pruning, so only the columns needed by a query are read from storage. This significantly reduces I/O and improves performance for analytical queries that access a subset of columns. Predicate pushdown further optimizes by filtering data at the source. Other formats like Avro or JSON are row-based and cannot provide these benefits.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable Azure Data Lake Storage Gen2 hierarchical namespace and organize files into folders by date.

    Why it's wrong here

    Hierarchical namespace improves directory operations and atomic renames, but it does not reduce the amount of data read during a query. Partitioning by date can help with partition pruning, but the scenario states that entire partitions are already scanned, and the issue is column pruning, not partition elimination. Folder organization alone does not optimize column-level reads.

  • ✓

    Use Parquet format with column pruning and predicate pushdown enabled in the query engine.

    Why this is correct

    Parquet is a columnar format that allows query engines to read only the columns referenced in a query, drastically reducing I/O. Predicate pushdown further filters data at the storage layer. In Azure Synapse Analytics and Databricks, Parquet supports these optimizations natively. This directly addresses the need to minimize data read and improve performance for column-subset queries.

  • ✗

    Store the data in a row-based format such as Avro to enable faster full-row retrieval.

    Why it's wrong here

    Avro is a row-based format, which means all columns in a row are read together. When queries need only a subset of columns, row-based formats force reading unnecessary data, increasing I/O and reducing performance. For analytical workloads with column pruning, a columnar format like Parquet is preferred. Avro is better suited for write-heavy or streaming scenarios where entire records are processed.

  • ✗

    Convert the data to JSON and compress it with GZip to reduce storage size.

    Why it's wrong here

    JSON is a row-based, text format that does not support column pruning. While GZip compression reduces storage footprint, it does not reduce the amount of data read for column-subset queries because the entire compressed file must be decompressed to access any part. This approach adds CPU overhead and fails to optimize query performance for analytical workloads.

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.