DP-900 Describe an analytics workload on Azure Practice Question
Which TWO Azure services can be used to store semi-structured data like JSON or Parquet files for analytics? (Choose two.)
⚠ Common exam trap
A common mix-up: candidates confuse Azure Synapse Analytics (a query service) with a storage service, or think Azure Cosmos DB is suitable for storing large Parquet files for analytics, when it is actually a transactional NoSQL database with a different cost and performance profile.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Azure Blob Storage
Azure Blob Storage (B) is correct because it provides massively scalable object storage that can hold semi-structured files such as JSON and Parquet in containers, and it is commonly used as a landing/staging area for analytics workloads. Azure Data Lake Storage Gen2 (D) is correct because it builds on Blob Storage with a hierarchical namespace, POSIX-like ACLs, and optimized performance for big-data analytics, making it the standard store for JSON, Parquet, and other semi-structured files consumed by engines like Synapse Spark and Databricks. Azure Synapse Analytics (A) is an analytics service with SQL and Spark pools rather than a primary storage service for JSON/Parquet files, Azure SQL Database (C) is a relational database management system designed for structured tabular data, and Azure Cosmos DB (E) is a globally distributed NoSQL database for operational document/key-value workloads, not a file-based analytics data store.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Azure Synapse Analytics
Why it's wrong here
Azure Synapse Analytics is an enterprise analytics service that unifies data warehousing and big data processing, but it is not a storage service itself. It can query semi-structured files stored in external locations via external tables or PolyBase, but the actual enduring data remains in underlying storage like Blob Storage or Data Lake Storage Gen2. Consequently, Synapse acts as a processing and analysis engine, not as one of the primary storage services for semi-structured data.
- ✓
Azure Blob Storage
Why this is correct
Azure Blob Storage is Microsoft's object storage service, ideal for storing massive amounts of unstructured and semi-structured data as blobs. It accepts any file type—JSON, Parquet, CSV, Avro—without a predefined schema, making it a correct choice for semi-structured data. Blob Storage provides tiered storage, lifecycle management, and high durability, and it serves as the foundation for Azure Data Lake Storage Gen2's hierarchical namespace.
- ✗
Azure SQL Database
Why it's wrong here
Azure SQL Database is a managed relational database service that stores data in tables with a defined schema, requiring entities to conform to precise column types and constraints. Although it supports JSON functions for parsing and querying JSON fragments, the JSON is still placed inside relational columns, and the service is not designed for storing raw semi-structured files or documents. Thus, it is not a general-purpose storage service for semi-structured data.
- ✓
Azure Data Lake Storage Gen2
Why this is correct
Azure Data Lake Storage Gen2 is a data lake solution built on Azure Blob Storage that adds a hierarchical namespace and POSIX-like access control lists, enabling efficient organization of files and directories. It supports all file formats, including semi-structured data such as JSON, Parquet, and ORC, while providing petabyte-scale storage and integrated analytics. This makes it a premier storage service for semi-structured datasets used in big data workloads.
- ✗
Azure Cosmos DB
Why it's wrong here
Azure Cosmos DB is a globally distributed, multi-model NoSQL database service that natively stores JSON documents, key-value records, graphs, and column-family data with a flexible schema. While it absolutely handles semi-structured data, it is designed as a transactional, queryable database with indexing, consistency levels, and SLAs, not as a raw file or object storage service. In the context of this question, Cosmos DB belongs to the database category, not the core storage solutions like Blob Storage and Data Lake Storage Gen2.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
Learn chapter
Data Formats: JSON, CSV, Parquet, and Avro
Key term
Service
A service is a software component or system that performs a specific function and is available to be used by other programs or users over a network.
Key term
Data Lake Storage Gen2
Data Lake Storage Gen2 is a cloud-based storage service that combines a scalable data lake with enterprise-grade file system capabilities for big data analytics.
About these practice questions
One of 851 original DP-900 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DP-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-900 exam.