Courseiva

DP-203 Practice Question: Secure, monitor, and optimize data storage and data processing

You have an Azure Data Lake Storage Gen2 account that contains a container named raw. The container has a folder hierarchy with millions of small files. You need to optimize read performance for an Azure Databricks job that reads these files. You also need to minimize storage costs. What should you do?

⚠ Common exam trap

The trap here is focusing on cluster scaling or storage tier when the real issue is the small file problem.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Compact the small files into larger files using a tool like Azure Data Factory or Databricks, and store them in a partitioned folder structure.

The presence of millions of small files creates significant overhead in listing, opening, and reading each file, which degrades performance and increases costs. Compacting these files into larger, partitioned files reduces the number of files and the amount of data scanned, improving read performance and lowering storage costs. The other options either do not address the root cause or increase costs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Move the files to a premium storage tier to reduce latency.

    Why it's wrong here

    Premium storage tiers can reduce latency for hot data, but they significantly increase storage costs, which conflicts with the requirement to minimize costs. Additionally, premium tiers do not address the small file overhead, which is the primary performance issue. Compacting files is a more cost-effective and direct solution.

  • ✓

    Compact the small files into larger files using a tool like Azure Data Factory or Databricks, and store them in a partitioned folder structure.

    Why this is correct

    Compacting many small files into fewer large files reduces the overhead of opening and reading each file, significantly improving read performance. Partitioning the data by a common filter column further reduces the amount of data scanned. This approach also lowers storage costs by reducing metadata overhead and improving compression efficiency, addressing both performance and cost goals.

  • ✗

    Increase the number of partitions in the Databricks cluster to parallelize reads across more nodes.

    Why it's wrong here

    Adding more nodes or partitions can improve parallelism, but with millions of small files, the bottleneck is often the per-file overhead and metadata operations. Simply adding resources may not solve the underlying issue and can increase costs without addressing the small file problem. The more effective solution is to compact files first.

  • ✗

    Enable hierarchical namespace and use Azure Blob Storage APIs to read the files.

    Why it's wrong here

    Hierarchical namespace is already required for Data Lake Storage Gen2 and enables directory operations, but it does not directly optimize read performance for millions of small files. Using Blob Storage APIs may bypass some optimizations and is not a performance tuning step. The core issue is file size and count, not the namespace or API choice.

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.