Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is designing an ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver streaming records into an Amazon S3 bucket. The records arrive as JSON, and downstream consumers require Parquet with a stable schema. The engineer must configure the Firehose delivery stream so records are converted to Parquet before landing in S3. (Choose two.)

⚠ Common exam trap

The trap here is assuming a Lambda transform or a file extension change can produce Parquet, when Firehose requires its built-in record format conversion with a Glue table schema reference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable record format conversion on the Firehose delivery stream and select Apache Parquet as the output format.

Firehose record format conversion converts JSON to Parquet in flight, and it requires a schema reference. The supported schema source is an AWS Glue table in the Data Catalog that describes the incoming JSON structure. Enabling conversion with Apache Parquet output and pointing to a matching Glue table together produce Parquet files in S3 for downstream analytics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable Amazon S3 Object Lambda access points on the destination bucket to convert objects to Parquet on read.

    Why it's wrong here

    S3 Object Lambda transforms objects when they are retrieved, not when they are written. It could theoretically present Parquet on read, but it does not satisfy the requirement that records land in S3 as Parquet for downstream consumers, and it adds read-time latency and complexity. Firehose's native conversion is the correct approach.

  • ✓

    Enable record format conversion on the Firehose delivery stream and select Apache Parquet as the output format.

    Why this is correct

    Firehose supports record format conversion from JSON to Parquet. Enabling it and selecting Apache Parquet as the output format instructs Firehose to convert records in flight before writing to S3. Without this setting, Firehose writes the raw JSON payload, so downstream consumers would not receive Parquet as required.

  • ✓

    Attach an AWS Glue table as the schema reference for the record format conversion configuration.

    Why this is correct

    Record format conversion in Firehose requires a schema so the service knows how to map JSON fields to Parquet columns. Firehose uses an AWS Glue table in the Data Catalog as that schema reference. The table must match the incoming JSON structure; otherwise conversion fails or produces nulls, so this configuration step is mandatory alongside enabling conversion.

  • ✗

    Configure an AWS Lambda function as the delivery stream's transformation to rewrite each record into Parquet bytes.

    Why it's wrong here

    A Lambda transform can modify records, but the Lambda runtime does not natively emit Parquet, and Firehose buffers records before writing. Writing Parquet per record is inefficient and does not produce the columnar files Firehose format conversion creates. Firehose's built-in record format conversion is the supported mechanism, so a Lambda rewrite is unnecessary and error-prone here.

  • ✗

    Set the S3 destination prefix to include a .parquet file extension so Firehose writes columnar files.

    Why it's wrong here

    The S3 prefix controls folder and key naming, not file encoding. Changing the prefix to end in .parquet does not convert JSON payloads into Parquet; it only changes the object key. The data would still be JSON bytes with a misleading extension, so downstream Parquet readers would fail to parse the files.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.