Courseiva
mediumMultiple Choice

DP-203 Practice Question: Design a near-real-time data processing solution…

You need to design a near-real-time data processing solution that ingests IoT telemetry data from millions of devices. The data must be aggregated per minute and stored in Azure Cosmos DB for low-latency queries. Which Azure service combination should you use?

⚠ Common exam trap

The trap here is that candidates often over-engineer the solution by adding a big-data processing layer (like HDInsight or Databricks) when a simpler, fully managed stream analytics service (Azure Stream Analytics) is the correct choice for fixed-window aggregation and direct Cosmos DB output.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Azure Event Hubs -> Azure Stream Analytics -> Azure Cosmos DB

Azure Stream Analytics provides native, low-latency windowed aggregation (e.g., TumblingWindow for per-minute aggregates) directly on data ingested from Event Hubs, and it has a built-in output sink to Azure Cosmos DB. This combination meets the near-real-time requirement without needing an intermediate compute or storage layer, minimizing end-to-end latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Azure Event Hubs -> Azure HDInsight (Kafka) -> Azure Cosmos DB

    Why it's wrong here

    HDInsight Kafka is a batch-oriented cluster provisioned for long-running analytics jobs, adding cluster management latency and cost rather than the continuous per-minute streaming aggregation required. HDInsight is the correct choice for large-scale batch processing or Hadoop ecosystem workloads, not near-real-time event ingestion.

  • ✓

    Azure Event Hubs -> Azure Stream Analytics -> Azure Cosmos DB

    Why this is correct

    Event Hubs ingests millions of device telemetry streams at scale, Stream Analytics performs the per-minute tumbling-window aggregation, and Cosmos DB stores results for low-latency queries. This combination satisfies both the near-real-time ingestion requirement and the minute-level aggregation constraint.

  • ✗

    Azure IoT Hub -> Azure Databricks (Structured Streaming) -> Azure Cosmos DB

    Why it's wrong here

    Databricks streaming can work but is less straightforward.

  • ✗

    Azure Event Hubs -> Azure Data Factory -> Azure Cosmos DB

    Why it's wrong here

    Data Factory runs on scheduled pipeline triggers, not continuous streaming, so it cannot aggregate per-minute telemetry from millions of devices in near-real time. Data Factory is the correct ingestion tool for batch or scheduled movement between data stores, not for streaming event pipelines.

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.