Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is designing a serverless data ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using AWS Lambda before being written to S3. Which two steps are required to enable this transformation? (Select TWO.)

⚠ Common exam trap

Test-takers frequently confuse post-delivery transformations (using S3 event notifications) with in-stream transformations (using Firehose's built-in Lambda integration), leading them to select Option A instead of the correct Firehose-specific configuration.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure a Lambda function as a data transformation source in the Firehose delivery stream.

Option B is correct because Firehose supports Lambda-based data transformation by letting you specify a Lambda function as the processor in the delivery stream's transform configuration (via the console or the ProcessingConfiguration/TransformParameters in the API), which Firehose then invokes synchronously for each buffered batch. Option C is correct because the Lambda function must return the transformed records in the exact structure Firehose expects — a JSON object containing records with recordId, result (Ok, Dropped, or ProcessingFailed), and base64-encoded data — otherwise Firehose cannot continue delivery. Option A is wrong because S3 event notifications trigger actions on object creation and play no role in Firehose's inline transformation; Firehose itself invokes the Lambda function. Option D is wrong because subscribing Lambda to the Firehose CloudWatch Logs log group is not a configuration step for transformation and would not enable it. Option E is wrong because the Lambda function must return transformed data to Firehose, not write directly to S3; Firehose remains responsible for delivering the records to the destination bucket.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set up an S3 event notification to trigger the Lambda function on object creation.

    Why it's wrong here

    Firehose invokes the Lambda transform itself and passes records through the processing configuration; an S3 event notification fires only after objects land, which is too late to transform them before delivery. It is tempting because S3 event notifications commonly trigger Lambda, but that pattern suits post-upload processing, not in-flight Firehose transformation.

  • ✓

    Configure a Lambda function as a data transformation source in the Firehose delivery stream.

    Why this is correct

    Firehose invokes a Lambda function only when it is attached as the delivery stream's data transformation source, which enables the buffered records to be processed before delivery. Without this configuration, Firehose writes raw records straight to S3, so transformation never occurs.

  • ✓

    Ensure the Lambda function returns the transformed data in the format required by Firehose.

    Why this is correct

    Firehose requires the transformation Lambda to return records in its prescribed structure, with each record carrying the processed payload plus a result status, so Firehose knows whether to deliver, drop or retry it. Returning arbitrary output causes the delivery stream to treat records as failed.

  • ✗

    Subscribe the Lambda function to the CloudWatch Logs log group for the Firehose stream.

    Why it's wrong here

    CloudWatch Logs subscription filters stream log events to Lambda for log processing; Firehose transformation requires the Lambda specified in the stream's processing configuration instead. It is tempting because subscribing Lambda to logs is a real integration, but it would be correct only for reacting to emitted log data, not transforming Firehose records.

  • ✗

    Have the Lambda function write the transformed data directly to the S3 bucket.

    Why it's wrong here

    Firehose writes the delivered objects to S3; the Lambda transform must return transformed records to Firehose, which then delivers them. Having Lambda write directly to S3 bypasses the delivery stream and duplicates its role. It is tempting because Lambda writing to S3 is common, but that suits standalone pipelines, not Firehose-managed delivery.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.