Courseiva
mediumMultiple Choice

DP-203 Practice Question: You have an Azure Data Lake Storage Gen2 account…

You have an Azure Data Lake Storage Gen2 account that stores large volumes of parquet files. A reporting application frequently queries a specific subset of data filtered by a 'region' column. To minimize query latency and cost, which optimization should you implement?

⚠ Common exam trap

Candidates often confuse compression (Option C) with partitioning, thinking reducing file size alone minimizes I/O, but without partition pruning the engine still scans all files, negating the benefit.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Partition the data by region in the folder structure.

Partitioning the data by region in the folder structure (e.g., /region=NorthAmerica/...) enables Azure Data Lake Storage Gen2 and query engines like Azure Synapse or PolyBase to perform partition pruning. This skips scanning irrelevant files entirely, reducing I/O and query latency while lowering cost by minimizing data processed.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Partition the data by region in the folder structure.

    Why this is correct

    Partitioning parquet files by region lets the reporting application prune folders and read only the relevant region's data. This reduces bytes scanned per query, cutting both latency and cost compared with scanning the full dataset on every request.

  • ✗

    Create a clustered index on the region column.

    Why it's wrong here

    Clustered indexes are a relational database construct; parquet files in Data Lake Storage are not indexed by the storage engine, so no such index can be created or used. It is tempting because index seeks are the standard relational remedy for filtered queries, and would be correct on a dedicated SQL pool table.

  • ✗

    Compress the parquet files using gzip.

    Why it's wrong here

    Gzip reduces bytes read but every file is still decompressed and scanned, since compression cannot skip row groups by region value. It is tempting because it lowers storage and egress cost, and would be correct when the bottleneck is transfer volume rather than filter selectivity.

  • ✗

    Enable hierarchical namespace on the storage account.

    Why it's wrong here

    Hierarchical namespace enables directory-level operations and atomic renames; it does not prune data by column value, so the region filter still scans every parquet file. It is tempting because it improves analytics workload performance generally, and would be correct when the workload needs POSIX-style directory semantics.

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.