Courseiva

DP-203 Practice Question: Secure, monitor, and optimize data storage and data processing

Which THREE best practices should be followed when designing a data lake in Azure Data Lake Storage Gen2 for optimal performance?

⚠ Common exam trap

The trap is that candidates assume disabling the hierarchical namespace improves performance (it actually removes key optimizations) and that deeper folder hierarchies are better, when flat, date-partitioned layouts are the recommended design.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Parquet file format for analytics workloads.

Option C is correct because columnar formats like Parquet compress data efficiently and let analytics engines such as Spark, Synapse, and Databricks read only the needed columns, dramatically reducing I/O and improving query performance. Option D is correct because ADLS Gen2 (and the underlying Blob REST/ABFS APIs) handles simple, predictable names more efficiently; special characters can break path parsing and high-cardinality names (e.g., GUIDs or timestamps in every filename) defeat caching, listing, and partition pruning. Option E is correct because partitioning data by date (e.g., year=/month=/day=) aligns with how engines like Spark, Hive, and Synapse perform partition elimination, so queries scanning a single day skip all other partitions instead of reading the whole dataset. Option A is wrong because enabling the hierarchical namespace is what makes ADLS Gen2 a true data lake with directory-level operations, atomic renames, and better performance for analytics; disabling it reverts to flat Blob storage semantics. Option B is wrong because deep, heavily nested directory trees increase metadata and listing overhead and slow down operations; a shallow, well-partitioned structure is preferred.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Disable hierarchical namespace to improve performance.

    Why it's wrong here

    Hierarchical namespace is what enables Data Lake Storage Gen2 semantics, directory-level operations and atomic renames; disabling it reduces the account to flat Blob storage. It tempts because flat blob namespaces avoid namespace overhead, which can help pure Blob workloads without directory operations.

  • ✗

    Use a deep directory structure with many subfolders.

    Why it's wrong here

    Deep nested folders add metadata and rename overhead, degrading Data Lake Storage Gen2 throughput at scale. Flat hierarchies with few levels suit parallel analytics workloads; deep structures are chosen only for strict organisational partitioning, not performance.

  • ✓

    Use Parquet file format for analytics workloads.

    Why this is correct

    Parquet's columnar, compressed layout lets analytical engines read only the referenced columns and skip irrelevant row groups, sharply cutting I/O against the stem's optimal-performance requirement. Row-based formats force full-row deserialisation regardless of projection, so Parquet directly satisfies the constraint that analytics workloads over Data Lake Storage Gen2 must minimise scanned data.

  • ✓

    Use a naming convention that avoids special characters and high cardinality.

    Why this is correct

    Avoiding special characters and high cardinality in file and directory names prevents inefficient path parsing and excessive metadata operations in Azure Data Lake Storage Gen2. High-cardinality names fragment the namespace, degrading listing and query performance, so this convention directly satisfies the stem's requirement for optimal performance.

  • ✓

    Partition data by date to enable partition elimination.

    Why this is correct

    Partitioning by date aligns the physical folder hierarchy with the query predicate, so the engine prunes irrelevant directories and reads only the required date ranges. This directly satisfies the stem's performance goal by cutting the volume of data scanned during analytical reads in Data Lake Storage Gen2.

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.