DP-900 Practice Question: Describe considerations for working with non-relational data on Azure
You are designing a solution to store and analyze large volumes of streaming data from social media feeds. The data is semi-structured (JSON) and will be used for real-time dashboards. You need to choose a storage solution that can handle high-ingestion throughput and support querying with Azure Synapse Serverless SQL. Which storage option should you choose?
⚠ Common exam trap
DP-900 often tests the confusion between transactional and analytical storage — candidates may pick Cosmos DB for its JSON support, but it is not designed for Synapse Serverless SQL analytical querying at scale.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is the correct choice because it is optimized for big data analytics, supports hierarchical namespace, and can store semi-structured data like JSON at scale. It integrates natively with Azure Synapse Serverless SQL, allowing direct querying of files using T-SQL. Its high-ingestion throughput and support for various file formats make it ideal for streaming data landing and analysis.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Azure Table Storage
Why it's wrong here
Azure Table Storage is a NoSQL key-value store designed for high-volume, low-latency transactional access to semi-structured data. It does not provide a SQL query engine, columnar storage, or partitioning that supports large-scale analytical scans like Synapse Serverless SQL requires. As a result, it is not suitable for analytics on large volumes of data, which is why it is incorrect for this solution.
- ✓
Azure Data Lake Storage Gen2
Why this is correct
Azure Data Lake Storage Gen2 is a hierarchical file system built on Azure Blob Storage that stores data in open formats such as Parquet and ORC, enabling massive parallel ingestion. Synapse Serverless SQL can query files directly using the OPENROWSET function with predicate pushdown to the storage layer, making it both fast and cost-efficient for big data analytics. This alignment with the analytic workload makes it the correct choice.
- ✗
Azure Cosmos DB
Why it's wrong here
Cosmos DB is a globally distributed, multi-model NoSQL database optimized for single-digit millisecond reads and writes on transactional data, but it cannot be directly queried by Synapse Serverless SQL for ad hoc analytical queries over large datasets. Its analytical features require a separate analytical store and Synapse Link, adding complexity and not matching the simplicity of querying files. Therefore it is not ideal as a large-volume analytical store.
- ✗
Azure Cache for Redis
Why it's wrong here
Redis is an in-memory key-value store used primarily to cache frequently accessed data and reduce latency, but it does not persist data to disk by default, offers no SQL schema, and cannot be scanned by Synapse Serverless SQL for analytical purposes. It is meant for sub-millisecond data retrieval within an application, not for storing historical or bulk data for analytics. Hence it is unsuitable for this solution.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
Learn chapter
Azure Synapse Link for Cosmos DB
Key term
Structured data
Structured data is information that is organized in a predefined format, typically in rows and columns, making it easy to search, process, and analyze by computers.
Key term
Semi-structured data
Semi-structured data is information that has some organizational tags or markers but does not fit into a strict table format like a spreadsheet row and column.
About these practice questions
One of 851 original DP-900 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This DP-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-900 exam.