DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is designing a data ingestion pipeline for real-time clickstream data. Which TWO services can be used to ingest the data into Amazon Kinesis Data Streams?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Kinesis Producer Library (KPL)
Options B and D are correct. The Kinesis Producer Library (KPL) is a library for producers to send data to Kinesis Data Streams. AWS SDK can also be used directly. Option A is wrong because Amazon S3 is a storage service, not a producer. Option C is wrong because Kinesis Data Firehose is a downstream consumer or delivery service, not a producer. Option E is wrong because AWS Glue is an ETL service, not a producer.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Amazon S3
Why it's wrong here
Amazon S3 is object storage, not a producer that writes records into Kinesis Data Streams; ingestion requires the Kinesis Producer Library, SDK or agent. It is tempting because S3 commonly holds clickstream archives, but that is batch storage feeding analytics, not real-time stream ingestion.
- ✓
Kinesis Producer Library (KPL)
Why this is correct
The Kinesis Producer Library aggregates and batches records, then sends them to Kinesis Data Streams using the PutRecords API. It handles retries and throughput optimisation, satisfying the real-time clickstream ingestion requirement without custom serialisation or buffering code.
- ✗
Kinesis Data Firehose
Why it's wrong here
Kinesis Data Firehose delivers stream data to destinations such as S3, Redshift and Splunk; it consumes from Data Streams rather than ingesting into them. It is tempting because it handles streaming delivery, but that is egress, whereas the stem requires producers writing records into Data Streams.
- ✓
AWS SDK
Why this is correct
The AWS SDK exposes the Kinesis Data Streams PutRecord and PutRecords APIs directly, letting applications publish clickstream events programmatically. This satisfies the ingestion requirement when you need fine-grained control rather than the KPL's aggregation layer.
- ✗
AWS Glue
Why it's wrong here
AWS Glue is a serverless ETL and catalog service that reads and transforms data; it does not act as a producer writing records into Kinesis Data Streams. It is tempting because Glue jobs commonly process streaming sources, but that is consumption and transformation, not ingestion into the stream.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.