Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A data scientist is building a recommendation system using Apache Spark for feature engineering. They need to process streaming user click data in real-time before feeding into the model. Which tool should they use for the streaming data ingestion?

⚠ Common exam trap

CompTIA often tests the distinction between storage, orchestration, and streaming tools, and the trap here is that candidates confuse batch-oriented tools like S3 or Airflow with real-time streaming ingestion, overlooking Kafka's role as a dedicated event streaming platform.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apache Kafka

Apache Kafka is the correct choice because it is a distributed streaming platform designed for high-throughput, fault-tolerant, real-time data ingestion. It acts as a durable message broker that can ingest streaming click data and make it available for Spark Structured Streaming to process in micro-batches or continuous processing mode, which is essential for real-time feature engineering in a recommendation system.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon S3

    Why it's wrong here

    Amazon S3 is object storage, not a streaming ingestion service; it cannot consume continuous click events in real time. It suits durable batch storage or a landing zone for processed data, but ingesting live streams requires a purpose-built streaming service such as Kinesis or Kafka.

  • ✓

    Apache Kafka

    Why this is correct

    Kafka provides a distributed, partitioned, replicated commit log that ingests high-throughput click streams with low latency and durable buffering, decoupling producers from Spark Structured Streaming consumers. This satisfies the requirement to process real-time streaming click data before feature engineering.

  • ✗

    Airflow

    Why it's wrong here

    Airflow orchestrates scheduled batch workflows; it triggers jobs on a timetable rather than continuously ingesting live click events. It suits coordinating dependent batch pipelines, not real-time streaming ingestion, which needs a persistent consumer such as Kinesis or Kafka feeding Spark.

  • ✗

    Snowflake

    Why it's wrong here

    Snowflake is a cloud data warehouse for storing and querying structured data, not a streaming ingestion service. It suits analytical SQL workloads over loaded datasets, but real-time click ingestion requires a streaming platform such as Kinesis or Kafka delivering events to Spark.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.