DEA-C01 Data Ingestion and Transformation Practice Question
A company uses Kinesis Data Firehose to deliver streaming data to S3. They need to transform the data by adding a timestamp and removing sensitive fields. Which TWO approaches can achieve this?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS Glue ETL to process data after delivery to S3
Option E is correct because Kinesis Data Firehose natively supports Lambda-based data transformation: you attach a Lambda function to the delivery stream, and Firehose invokes it synchronously on each record (or batch) before delivery, allowing the function to add a timestamp and strip sensitive fields. Option C is correct because AWS Glue ETL can process the data after it lands in S3, using Spark-based jobs to enrich records with timestamps and drop sensitive columns, which is a valid post-delivery transformation approach. Option A is not correct because Kinesis Data Analytics is for real-time SQL/Flink analytics on streams, not for modifying records delivered by Firehose. Option B is not correct because S3 Select only filters and projects data at rest using SQL on individual objects; it cannot add timestamps or rewrite/remove fields. Option D is not correct because Redshift Spectrum queries data in S3 for analytics and does not transform or modify the delivered objects.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Kinesis Data Analytics to transform the stream
Why it's wrong here
Kinesis Data Analytics runs SQL or Flink over the stream, but Firehose cannot consume its output directly as an inline transform; the stem requires transformation within the delivery pipeline. A Lambda processor invoked by Firehose performs the timestamp addition and field removal. Data Analytics suits continuous stream analytics, not Firehose record rewriting.
- ✗
Use S3 Select to transform data at rest
Why it's wrong here
S3 Select filters and projects data during retrieval from S3; it cannot add a timestamp or rewrite objects as Firehose delivers them. In-flight transformation needs a Firehose Lambda processor. S3 Select is correct when querying a subset of columns from stored objects to reduce data transfer.
- ✓
Use AWS Glue ETL to process data after delivery to S3
Why this is correct
Running AWS Glue ETL over the delivered S3 objects lets you add the timestamp and drop sensitive fields after Firehose lands the data. This satisfies the transformation requirement using a managed, serverless job rather than altering the delivery stream itself.
- ✗
Use Amazon Redshift Spectrum to transform data
Why it's wrong here
Redshift Spectrum queries data already in S3 or Redshift; it cannot intercept records inside a Firehose delivery stream. Transformation must occur in-flight via a Lambda function invoked by Firehose. Spectrum is correct for ad-hoc SQL analytics over large S3 datasets, not stream processing.
- ✓
Configure a Lambda function as a data transformation in Firehose
Why this is correct
Firehose invokes a Lambda function synchronously on each buffered batch, letting custom code add timestamps and strip sensitive fields before delivery. This satisfies the in-flight transformation requirement without managing consumers, since Firehose handles buffering, retries and S3 delivery natively.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.