Courseiva

DP-900 Describe an analytics workload on Azure Practice Question

Which TWO Azure services can be used to store semi-structured data like JSON or Parquet files for analytics? (Choose two.)

⚠ Common exam trap

A common mix-up: candidates confuse Azure Synapse Analytics (a query service) with a storage service, or think Azure Cosmos DB is suitable for storing large Parquet files for analytics, when it is actually a transactional NoSQL database with a different cost and performance profile.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Azure Blob Storage

Azure Blob Storage (B) is correct because it provides massively scalable object storage that can hold semi-structured files such as JSON and Parquet in containers, and it is commonly used as a landing/staging area for analytics workloads. Azure Data Lake Storage Gen2 (D) is correct because it builds on Blob Storage with a hierarchical namespace, POSIX-like ACLs, and optimized performance for big-data analytics, making it the standard store for JSON, Parquet, and other semi-structured files consumed by engines like Synapse Spark and Databricks. Azure Synapse Analytics (A) is an analytics service with SQL and Spark pools rather than a primary storage service for JSON/Parquet files, Azure SQL Database (C) is a relational database management system designed for structured tabular data, and Azure Cosmos DB (E) is a globally distributed NoSQL database for operational document/key-value workloads, not a file-based analytics data store.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Azure Synapse Analytics

    Why it's wrong here

    Azure Synapse Analytics is an enterprise analytics service that unifies data warehousing and big data processing, but it is not a storage service itself. It can query semi-structured files stored in external locations via external tables or PolyBase, but the actual enduring data remains in underlying storage like Blob Storage or Data Lake Storage Gen2. Consequently, Synapse acts as a processing and analysis engine, not as one of the primary storage services for semi-structured data.

  • ✓

    Azure Blob Storage

    Why this is correct

    Azure Blob Storage is Microsoft's object storage service, ideal for storing massive amounts of unstructured and semi-structured data as blobs. It accepts any file type—JSON, Parquet, CSV, Avro—without a predefined schema, making it a correct choice for semi-structured data. Blob Storage provides tiered storage, lifecycle management, and high durability, and it serves as the foundation for Azure Data Lake Storage Gen2's hierarchical namespace.

  • ✗

    Azure SQL Database

    Why it's wrong here

    Azure SQL Database is a managed relational database service that stores data in tables with a defined schema, requiring entities to conform to precise column types and constraints. Although it supports JSON functions for parsing and querying JSON fragments, the JSON is still placed inside relational columns, and the service is not designed for storing raw semi-structured files or documents. Thus, it is not a general-purpose storage service for semi-structured data.

  • ✓

    Azure Data Lake Storage Gen2

    Why this is correct

    Azure Data Lake Storage Gen2 is a data lake solution built on Azure Blob Storage that adds a hierarchical namespace and POSIX-like access control lists, enabling efficient organization of files and directories. It supports all file formats, including semi-structured data such as JSON, Parquet, and ORC, while providing petabyte-scale storage and integrated analytics. This makes it a premier storage service for semi-structured datasets used in big data workloads.

  • ✗

    Azure Cosmos DB

    Why it's wrong here

    Azure Cosmos DB is a globally distributed, multi-model NoSQL database service that natively stores JSON documents, key-value records, graphs, and column-family data with a flexible schema. While it absolutely handles semi-structured data, it is designed as a transactional, queryable database with indexing, consistency levels, and SLAs, not as a raw file or object storage service. In the context of this question, Cosmos DB belongs to the database category, not the core storage solutions like Blob Storage and Data Lake Storage Gen2.

Quick reference

Azure Blob Storage Tier Comparison

TierStorage CostRetrieval CostLatencyUse Case
HotHighestLowestImmediateActive data, frequent reads
CoolLowerHigherImmediateData accessed < once / month
ColdLower stillHigherImmediateData accessed < once / quarter
ArchiveLowestHighest + rehydration delayHoursLong-term compliance retention

About these practice questions

One of 851 original DP-900 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-900 exam.