Courseiva

Streaming Data Ingestion into Amazon S3

Which TWO AWS services can be used to ingest streaming data into Amazon S3? (Choose two.)

Quick Answer

The answer is Amazon Kinesis Data Firehose and Amazon MSK (Managed Streaming for Apache Kafka). Kinesis Data Firehose is purpose-built for streaming data ingestion into Amazon S3, as it can directly load real-time data streams into S3 without requiring custom code, while Amazon MSK can integrate with Kafka Connect to reliably sink streaming data into S3 using a connector. On the AWS Certified Data Engineer Associate DEA-C01 exam, this question tests your understanding of which services handle continuous, low-latency data ingestion versus batch or offline transfer methods. A common trap is confusing S3 Transfer Acceleration, which only speeds up existing uploads, or Snowball, which is for offline bulk data migration, with true streaming ingestion. Remember the mnemonic “Firehose flows, MSK connects” to recall that Kinesis Data Firehose delivers directly and MSK uses connectors for S3 sinks.

⚠ Common exam trap

Candidates often confuse Amazon S3 Transfer Acceleration (a speed optimization for existing uploads) with a streaming ingestion service, or they mistakenly think EBS or Snowball can handle real-time streaming data when they are designed for persistent block storage and offline bulk transfer, respectively.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Amazon Managed Streaming for Apache Kafka (Amazon MSK)

Amazon Kinesis Data Firehose is the easiest way to reliably load streaming data into Amazon S3. It can capture, transform, and deliver streaming data to S3 destinations in near real-time with no code required. Amazon MSK (Managed Streaming for Apache Kafka) can also ingest streaming data into S3 by using Kafka Connect with an S3 sink connector, which writes data from Kafka topics directly to S3.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon S3 Transfer Acceleration

    Why it's wrong here

    Transfer Acceleration speeds up transfers of existing objects to S3 over long distances using edge locations; it does not ingest streaming data. It is tempting because it genuinely accelerates uploads, and would be correct for large batch migrations or geographically distant clients uploading files.

  • ✓

    Amazon Managed Streaming for Apache Kafka (Amazon MSK)

    Why this is correct

    Amazon MSK runs Apache Kafka clusters whose topics hold streaming records; consumers, or Kafka Connect's S3 sink connector, read those topics and write objects into Amazon S3. This satisfies the stem's requirement for a service that ingests streaming data into S3.

  • ✓

    Amazon Kinesis Data Firehose

    Why this is correct

    Kinesis Data Firehose is a fully managed delivery stream that buffers incoming records and writes them to Amazon S3 as objects, optionally transforming them en route. This directly satisfies the stem's requirement for a service that ingests streaming data into S3.

  • ✗

    Amazon Elastic Block Store (Amazon EBS)

    Why it's wrong here

    EBS provides block storage volumes attached to EC2 instances; it is not an ingestion service and does not write to S3. It is tempting because EC2 can run streaming producers, but EBS itself would be correct only for instance-local persistent disk storage.

  • ✗

    AWS Snowball

    Why it's wrong here

    Snowball is a physical appliance for shipping bulk data offline into S3; it cannot ingest continuous streams. It is tempting because it does move data into S3, and would be correct for migrating terabytes of existing on-premises data where network transfer is impractical.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on DEA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. Which TWO services can be used to ingest streaming data into Amazon S3? (Choose two.)

medium
  • A.Amazon Athena
  • B.AWS Glue
  • ✓ C.Amazon Kinesis Data Streams
  • D.AWS Database Migration Service (DMS)
  • ✓ E.Amazon Kinesis Data Firehose

Why C: Amazon Kinesis Data Streams (C) is correct because it is a real-time streaming service that can deliver records to Amazon S3 as a destination via Kinesis Data Firehose or custom consumers, enabling continuous ingestion of streaming data into S3. Amazon Kinesis Data Firehose (E) is correct because it is a fully managed service designed to reliably load streaming data directly into Amazon S3 (as well as Redshift, OpenSearch, and Splunk), making it a primary AWS solution for streaming ingestion into S3. Amazon Athena (A) is incorrect because it is an interactive query service for analyzing data already in S3, not a streaming ingestion mechanism. AWS Glue (B) is incorrect because it is a serverless ETL and data catalog service that processes batch or micro-batch jobs, not a native streaming ingestion service into S3. AWS Database Migration Service (D) is incorrect because it migrates and replicates databases, which is not the same as ingesting streaming data into S3.

Variation 2. Which TWO AWS services can be used to ingest streaming data into Amazon S3 with minimal code? (Choose two.)

medium
  • A.AWS Lambda
  • ✓ B.Amazon Kinesis Data Firehose
  • ✓ C.Amazon Managed Streaming for Apache Kafka (MSK) with S3 sink connector
  • D.AWS Database Migration Service (DMS)
  • E.AWS DataSync

Why B: Amazon Kinesis Data Firehose (B) is correct because it is a fully managed service that can deliver streaming data directly into Amazon S3 with essentially no code — you simply configure a delivery stream with S3 as the destination, and Firehose handles buffering, compression, encryption, and retries. Amazon MSK with an S3 sink connector (C) is also correct because the managed Kafka Connect S3 sink connector ingests streaming records from Kafka topics into S3 with only configuration rather than custom application code. AWS Lambda (A) is not the best fit here because it requires you to write and maintain function code to read from a stream and write objects to S3. AWS DMS (D) is designed for migrating and replicating databases, not for general streaming ingestion into S3. AWS DataSync (E) is a data-transfer service for moving files between on-premises storage and AWS, not for ingesting streaming data.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.