Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is designing a data ingestion pipeline for real-time clickstream data. Which TWO services can be used to ingest the data into Amazon Kinesis Data Streams?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Kinesis Producer Library (KPL)

Options B and D are correct. The Kinesis Producer Library (KPL) is a library for producers to send data to Kinesis Data Streams. AWS SDK can also be used directly. Option A is wrong because Amazon S3 is a storage service, not a producer. Option C is wrong because Kinesis Data Firehose is a downstream consumer or delivery service, not a producer. Option E is wrong because AWS Glue is an ETL service, not a producer.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon S3

    Why it's wrong here

    Amazon S3 is object storage, not a producer that writes records into Kinesis Data Streams; ingestion requires the Kinesis Producer Library, SDK or agent. It is tempting because S3 commonly holds clickstream archives, but that is batch storage feeding analytics, not real-time stream ingestion.

  • ✓

    Kinesis Producer Library (KPL)

    Why this is correct

    The Kinesis Producer Library aggregates and batches records, then sends them to Kinesis Data Streams using the PutRecords API. It handles retries and throughput optimisation, satisfying the real-time clickstream ingestion requirement without custom serialisation or buffering code.

  • ✗

    Kinesis Data Firehose

    Why it's wrong here

    Kinesis Data Firehose delivers stream data to destinations such as S3, Redshift and Splunk; it consumes from Data Streams rather than ingesting into them. It is tempting because it handles streaming delivery, but that is egress, whereas the stem requires producers writing records into Data Streams.

  • ✓

    AWS SDK

    Why this is correct

    The AWS SDK exposes the Kinesis Data Streams PutRecord and PutRecords APIs directly, letting applications publish clickstream events programmatically. This satisfies the ingestion requirement when you need fine-grained control rather than the KPL's aggregation layer.

  • ✗

    AWS Glue

    Why it's wrong here

    AWS Glue is a serverless ETL and catalog service that reads and transforms data; it does not act as a producer writing records into Kinesis Data Streams. It is tempting because Glue jobs commonly process streaming sources, but that is consumption and transformation, not ingestion into the stream.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.