Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A company is ingesting streaming data from social media feeds using Amazon Kinesis Data Streams. The data volume peaks at 10,000 records per second, and each record is up to 1 KB. The company needs to archive the raw data in Amazon S3 in near real-time and also make it available for real-time analytics using Amazon Kinesis Data Analytics. What is the MOST efficient architecture to meet these requirements?

⚠ Common exam trap

The trap here is that candidates often overcomplicate the architecture by adding unnecessary Lambda functions or duplicate streams, not realizing that Kinesis Data Firehose and Kinesis Data Analytics can both consume from the same Kinesis Data Stream natively.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Kinesis Data Streams as the ingestion point. Use Kinesis Data Firehose to read from the stream and write to S3. Use Kinesis Data Analytics to read directly from the same stream.

Kinesis Data Streams can serve as a single ingestion point, with Kinesis Data Firehose reading from the stream to deliver data to S3 (with optional transformation) and Kinesis Data Analytics reading directly from the same stream for real-time analytics. This avoids unnecessary duplication of streams or Lambda-based processing, which would add latency and complexity. The architecture is the most efficient as it leverages native integrations without intermediate compute.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Kinesis Data Streams as the ingestion point. Use Kinesis Data Firehose to read from the stream, convert to Parquet, and write to S3. Use a Lambda function to send data to Kinesis Data Analytics.

    Why it's wrong here

    Firehose converts records to Parquet, which changes the raw payload, so the archived data is no longer the raw stream the requirement specifies. It is tempting because Firehose is the managed S3 delivery path, and Parquet is correct when the archive should be query-optimised rather than raw.

  • ✗

    Use Kinesis Data Streams as the ingestion point. Use a Lambda function to read from the stream, write to S3, and send data to Kinesis Data Analytics.

    Why it's wrong here

    A Lambda consumer must poll the stream, run its own S3 batching logic, and separately fan out to Kinesis Data Analytics, adding custom code and per-invocation overhead at 10,000 records per second. It is tempting because Lambda suits lightweight stream processing, but managed delivery to S3 is Kinesis Data Firehose's purpose.

  • ✗

    Use two Kinesis Data Streams: one for S3 delivery and one for Kinesis Data Analytics.

    Why it's wrong here

    Two streams duplicate ingestion, cost, and shard management, and the producer must write each record twice; a single stream supports multiple consumers reading independently. It is tempting because separating pipelines appears to isolate workloads, but that isolation is achieved with consumer applications, not duplicate streams.

  • ✓

    Use Kinesis Data Streams as the ingestion point. Use Kinesis Data Firehose to read from the stream and write to S3. Use Kinesis Data Analytics to read directly from the same stream.

    Why this is correct

    Kinesis Data Firehose consumes directly from the stream and buffers records before delivering them to Amazon S3, satisfying the near real-time archival requirement without custom consumer code. Kinesis Data Analytics reads the same stream independently, so both consumers run in parallel. This decouples archival from analytics, avoiding duplicate ingestion and extra compute.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.