Courseiva
Data Ingestion and TransformationhardMultiple ChoiceObjective-mapped

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is designing a data ingestion pipeline for clickstream data from a mobile app. The data volume varies, with occasional spikes up to 10 MB/s. The pipeline must persist the raw data in Amazon S3 and make it available for near-real-time analytics via Amazon Athena. Which combination of services minimizes cost and operational overhead?

⚠ Common exam trap

A common mix-up: candidates choose Amazon Kinesis Data Streams with Lambda (Option C) because they think it provides more control, but they overlook Lambda's concurrency limits and the operational burden of managing stream shards, making Firehose the simpler and cheaper choice for raw data ingestion to S3.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Amazon Kinesis Data Firehose with direct delivery to Amazon S3, then Amazon Athena

Amazon Kinesis Data Firehose is the most cost-effective and low-overhead solution for ingesting variable-volume clickstream data (up to 10 MB/s) into Amazon S3 because it is a fully managed service that automatically scales, buffers, and compresses data before delivery. It integrates directly with S3 without requiring custom code or infrastructure management, and the data is immediately queryable by Amazon Athena with no additional transformation steps.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Amazon Kinesis Data Streams with Amazon Kinesis Data Analytics, then Amazon S3

    Why it's wrong here

    Overkill for simple persistence; adds unnecessary cost.

  • Amazon SQS with an Auto Scaling group of EC2 instances writing to Amazon S3

    Why it's wrong here

    Requires managing EC2 and scaling policies.

  • Amazon Kinesis Data Streams with AWS Lambda for transformation, then Amazon S3

    Why it's wrong here

    Kinesis Data Streams has per-shard costs and Lambda adds complexity.

  • Amazon Kinesis Data Firehose with direct delivery to Amazon S3, then Amazon Athena

    Why this is correct

    Firehose is fully managed, scales automatically, and delivers to S3.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,711 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.