Courseiva

DP-203 Design and implement data storage Practice Question

You need to design a data storage solution for a batch processing pipeline that processes petabytes of data daily. The data is stored in Parquet format and must be accessible by both Azure Databricks and Azure Synapse Analytics. Which storage solution should you recommend?

⚠ Common exam trap

Many candidates confuse Azure Blob Storage (flat namespace) with ADLS Gen2 (hierarchical namespace), assuming both are equivalent for big data analytics, but the hierarchical namespace is a critical differentiator for performance at petabyte scale in batch pipelines.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 (ADLS Gen2) is the correct choice because it combines a hierarchical namespace with Azure Blob Storage's scalable object storage, providing native POSIX-like access control and high throughput for petabyte-scale batch processing. Both Azure Databricks and Azure Synapse Analytics have optimized connectors for ADLS Gen2 that leverage the hierarchical namespace for efficient partition pruning and file listing, which is critical for Parquet-based analytics at this scale.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Azure Data Lake Storage Gen2

    Why this is correct

    Azure Data Lake Storage Gen2 provides a hierarchical namespace over Blob Storage, which is the specific capability enabling efficient directory-level operations across petabytes. Its POSIX-like ACLs and native ABFS driver satisfy the stem's requirement that both Azure Databricks and Azure Synapse Analytics access the same Parquet data without copying or conversion.

  • ✗

    Azure Files

    Why it's wrong here

    Azure Files exposes SMB and NFS shares with a 100 TiB limit and per-file overhead, so it cannot serve petabyte-scale Parquet to analytics engines efficiently. It is tempting because it provides managed file shares, and would be correct for lift-and-shift of on-premises SMB workloads or shared configuration files.

  • ✗

    Azure SQL Database

    Why it's wrong here

    Azure SQL Database caps at 4 TB per database and stores row-based relational tables, so it cannot hold petabytes of Parquet files or expose them as a filesystem. It is tempting because it is a fully managed relational engine, and would be correct for transactional OLTP workloads with structured schemas and ACID guarantees.

  • ✗

    Azure Blob Storage

    Why it's wrong here

    Blob Storage lacks a hierarchical namespace and POSIX semantics, so neither Azure Databricks nor Synapse Analytics gets the directory-level operations and atomic renames the pipeline needs. It is tempting because Blob Storage is cheap, massively scalable object storage, and would be correct for archival or serving static files to web applications.

Quick reference

Azure Blob Storage Tier Comparison

TierStorage CostRetrieval CostLatencyUse Case
HotHighestLowestImmediateActive data, frequent reads
CoolLowerHigherImmediateData accessed < once / month
ColdLower stillHigherImmediateData accessed < once / quarter
ArchiveLowestHighest + rehydration delayHoursLong-term compliance retention

Go deeper

Related to this question

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.