Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A data engineering team is designing a data pipeline to process streaming sensor data and feed it into an ML model for anomaly detection. Which THREE components are essential for this pipeline?

⚠ Common exam trap

CompTIA often tests the distinction between batch and streaming technologies, and the trap here is that candidates confuse Airflow's scheduling capability with real-time streaming orchestration, or assume Snowflake can act as a streaming sink when it is fundamentally a batch-oriented warehouse.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Amazon S3 as a data lake for storing raw sensor data

Amazon S3 is essential as a data lake for storing raw sensor data because it provides durable, scalable, and cost-effective object storage that can serve as a central repository for streaming data before and after processing. In a streaming pipeline, raw data must be persisted for reprocessing, historical analysis, and compliance, and S3's integration with Apache Spark and Kafka makes it a natural landing zone for sensor data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apache Airflow for scheduling recurring batch jobs

    Why it's wrong here

    Airflow is for orchestrating batch workflows, not for real-time streaming ingestion or processing.

  • ✓

    Amazon S3 as a data lake for storing raw sensor data

    Why this is correct

    S3 is a scalable object store that can serve as a data lake for raw sensor data, accessible for both streaming and batch processing.

  • ✗

    Snowflake as a real-time streaming destination

    Why it's wrong here

    Snowflake is a cloud data warehouse optimized for analytical queries on structured data, not designed for real-time streaming ingestion.

  • ✓

    Apache Kafka for ingesting streaming sensor data

    Why this is correct

    Kafka is a distributed streaming platform designed for high-throughput, fault-tolerant data ingestion.

  • ✓

    Apache Spark Structured Streaming for real-time processing

    Why this is correct

    Spark Structured Streaming allows real-time processing of streaming data with ML integration.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.