DP-900 Describe an analytics workload on Azure Practice Question
Your company is designing a big data analytics solution on Azure. The solution must ingest streaming data from IoT devices, store the data in its raw format, and then use a distributed processing engine to transform the data before loading it into a serving layer for reporting. Which TWO Azure services should you include in the design?
⚠ Common exam trap
Many exam-takers confuse Azure Event Hubs with Azure Blob Storage or Azure Data Factory for streaming ingestion, mistakenly thinking that any storage or ETL service can handle real-time IoT data, when in fact only a dedicated event ingestion service like Event Hubs provides the necessary throughput, partitioning, and protocol support for streaming workloads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Azure Event Hubs
Azure Event Hubs (B) is correct because it is Azure's managed, scalable event-ingestion service designed for high-throughput streaming telemetry from IoT devices, supporting millions of events per second and AMQP/Kafka protocols for real-time ingestion. Azure Databricks (C) is correct because it provides an Apache Spark-based distributed processing engine that can read the raw streamed data and run transformations at scale before writing to a serving layer. Together they satisfy the ingest-streaming-data and distributed-transformation requirements. Azure Blob Storage (A) is a storage service, not a streaming ingestion or distributed processing engine, so it does not fulfill the stated roles. Azure Data Factory (D) is an orchestration/ETL pipeline service rather than a streaming ingest or Spark processing engine. Azure Synapse Analytics (E) is primarily a serving/analytics warehouse layer, not the required streaming ingestion or distributed transformation engine.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Azure Blob Storage
Why it's wrong here
Azure Blob Storage provides massively scalable object storage for unstructured data, but it is a storage service rather than an event ingestion or stream processing platform. It cannot natively capture or deliver high-throughput streaming events with per-partition ordering and checkpointing; you'd need separate services to produce or consume data. It's typically used as a destination for processed results or as a data lake, not as the real-time ingestion layer for IoT analytics.
- ✓
Azure Event Hubs
Why this is correct
Azure Event Hubs is a fully managed, real-time data ingestion platform that can receive millions of events per second from devices using protocols like AMQP, HTTPS, and Apache Kafka. It provides a partitioned consumer model, configurable retention, and replay capability, making it the recommended front door for telemetry and IoT data. Once ingested, data can be routed to processing engines like Azure Stream Analytics or Databricks for transformation.
- ✓
Azure Databricks
Why this is correct
Azure Databricks is an Apache Spark-based analytics platform that provides distributed processing for large-scale data transformations, including streaming via Structured Streaming. It can read directly from Event Hubs, perform real-time transformations such as filtering, aggregating, and joining, and write results to storage or serving layers. While it is the compute engine for stream processing, it does not serve as the primary ingestion service, which is the role of Event Hubs.
- ✗
Azure Data Factory
Why it's wrong here
Azure Data Factory is a cloud ETL and data integration service that orchestrates and schedules data movement and transformations across pipelines. It is batch-oriented and trigger-based, rather than designed for sub-second, event-driven stream processing. While it can copy data continuously or on a schedule, it lacks the low-latency event handling and streaming semantics required for real-time IoT analytics.
- ✗
Azure Synapse Analytics
Why it's wrong here
Azure Synapse Analytics is an integrated analytics platform that unifies data warehousing, big data analytics, and data integration, but it primarily serves as a serving layer for querying and analyzing stored data. It does not provide native stream ingestion or sub-second stream processing; instead, it relies on other services to populate its SQL pools. As such, it is not the correct choice for the real-time ingestion and transformation stage.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
Learn chapter
Data Warehouse Distribution: Hash, Round-Robin, Replicated
Key term
Pipeline
A pipeline is an automated series of steps that takes code from development to production, ensuring quality and speed.
Key term
Blob storage
Blob storage is a cloud service for storing large amounts of unstructured data, such as text or binary data, like documents, images, and videos.
About these practice questions
One of 851 original DP-900 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DP-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-900 exam.