Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A company streams sensor data from IoT devices. The data arrives as JSON messages at high velocity. Which data pipeline architecture is BEST suited to handle this streaming data for near-real-time analytics?

⚠ Common exam trap

CompTIA often tests the distinction between batch and stream processing by presenting batch options that seem 'reliable' or 'traditional,' trapping candidates who overlook the explicit 'near-real-time' requirement in the question.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Stream processing using Apache Kafka and Spark Streaming.

Apache Kafka acts as a distributed, fault-tolerant ingestion layer that can handle high-velocity JSON messages, while Spark Streaming processes the data in micro-batches for near-real-time analytics. This combination provides the low-latency, scalable pipeline required for streaming IoT sensor data, unlike batch or single-node approaches.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Batch processing using Hadoop MapReduce every 24 hours.

    Why it's wrong here

    Hadoop MapReduce runs discrete 24-hour batch jobs over stored data, introducing latency that defeats near-real-time analytics on high-velocity sensor streams. MapReduce is the right choice for large-scale offline processing of accumulated data, not for continuous JSON message ingestion.

  • ✗

    Batch processing using nightly ETL jobs.

    Why it's wrong here

    Nightly ETL jobs process data in scheduled batches, so analytics results lag by up to 24 hours and cannot meet near-real-time requirements for high-velocity JSON sensor messages. Batch ETL suits periodic reporting on large historical datasets, not continuous streaming ingestion.

  • ✗

    Single-node database with periodic inserts.

    Why it's wrong here

    A single-node database with periodic inserts cannot sustain continuous high-velocity JSON ingestion, because writes are batched rather than processed as an unbounded event stream. It suits small-scale, low-frequency reporting where latency is tolerable. Near-real-time analytics requires a distributed streaming pipeline that ingests and processes each message as it arrives.

  • ✓

    Stream processing using Apache Kafka and Spark Streaming.

    Why this is correct

    Apache Kafka ingests the high-velocity JSON sensor messages durably as a distributed log, while Spark Streaming consumes those partitions and performs micro-batch analytics, satisfying the near-real-time requirement. Unlike batch pipelines, this architecture processes each message as it arrives rather than waiting for scheduled windows.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.