Courseiva
Data Ingestion and TransformationeasyMultiple ChoiceObjective-mapped

DEA-C01 Data Ingestion and Transformation Practice Question

A company is streaming clickstream data from a website into Amazon Kinesis Data Streams. The data must be transformed in near real-time and stored in Amazon S3 for analytics. Which AWS service should be used to transform the data as it is ingested?

⚠ Common exam trap

Test-takers frequently confuse AWS Glue's batch ETL capabilities with real-time streaming, or assume Lambda is always the best choice for stream processing, overlooking Kinesis Data Analytics' native support for continuous, stateful transformations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Amazon Kinesis Data Analytics

Amazon Kinesis Data Analytics is the correct choice because it can process and transform streaming data in near real-time using SQL or Apache Flink, and then output the transformed data to destinations like Amazon S3. This service is specifically designed for real-time stream processing, making it ideal for transforming clickstream data as it is ingested into Kinesis Data Streams.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • AWS Lambda (streaming function)

    Why it's wrong here

    AWS Lambda can process Kinesis records but is not ideal for high-throughput or complex transformations.

  • Amazon EMR (Spark Streaming)

    Why it's wrong here

    Amazon EMR with Spark Streaming is designed for large-scale, complex batch and stream processing jobs that require custom code and cluster management, but the scenario demands a lightweight, serverless transformation during ingestion without provisioning or managing infrastructure. It is tempting because Spark Streaming excels at sophisticated transformations on high-volume streams, and would be correct if the requirement involved complex analytics or machine learning on the data rather than simple, near real-time enrichment before landing in S3.

  • AWS Glue (ETL jobs)

    Why it's wrong here

    AWS Glue is primarily a batch ETL service, not designed for real-time stream processing.

  • Amazon Kinesis Data Analytics

    Why this is correct

    Amazon Kinesis Data Analytics can process and transform streaming data in real-time using SQL or Apache Flink.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,711 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.