DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is designing an ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver streaming records into an Amazon S3 bucket. The records arrive as JSON, and downstream consumers require Parquet with a stable schema. The engineer must configure the Firehose delivery stream so records are converted to Parquet before landing in S3. (Choose two.)
⚠ Common exam trap
The trap here is assuming a Lambda transform or a file extension change can produce Parquet, when Firehose requires its built-in record format conversion with a Glue table schema reference.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable record format conversion on the Firehose delivery stream and select Apache Parquet as the output format.
Firehose record format conversion converts JSON to Parquet in flight, and it requires a schema reference. The supported schema source is an AWS Glue table in the Data Catalog that describes the incoming JSON structure. Enabling conversion with Apache Parquet output and pointing to a matching Glue table together produce Parquet files in S3 for downstream analytics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable Amazon S3 Object Lambda access points on the destination bucket to convert objects to Parquet on read.
Why it's wrong here
S3 Object Lambda transforms objects when they are retrieved, not when they are written. It could theoretically present Parquet on read, but it does not satisfy the requirement that records land in S3 as Parquet for downstream consumers, and it adds read-time latency and complexity. Firehose's native conversion is the correct approach.
- ✓
Enable record format conversion on the Firehose delivery stream and select Apache Parquet as the output format.
Why this is correct
Firehose supports record format conversion from JSON to Parquet. Enabling it and selecting Apache Parquet as the output format instructs Firehose to convert records in flight before writing to S3. Without this setting, Firehose writes the raw JSON payload, so downstream consumers would not receive Parquet as required.
- ✓
Attach an AWS Glue table as the schema reference for the record format conversion configuration.
Why this is correct
Record format conversion in Firehose requires a schema so the service knows how to map JSON fields to Parquet columns. Firehose uses an AWS Glue table in the Data Catalog as that schema reference. The table must match the incoming JSON structure; otherwise conversion fails or produces nulls, so this configuration step is mandatory alongside enabling conversion.
- ✗
Configure an AWS Lambda function as the delivery stream's transformation to rewrite each record into Parquet bytes.
Why it's wrong here
A Lambda transform can modify records, but the Lambda runtime does not natively emit Parquet, and Firehose buffers records before writing. Writing Parquet per record is inefficient and does not produce the columnar files Firehose format conversion creates. Firehose's built-in record format conversion is the supported mechanism, so a Lambda rewrite is unnecessary and error-prone here.
- ✗
Set the S3 destination prefix to include a .parquet file extension so Firehose writes columnar files.
Why it's wrong here
The S3 prefix controls folder and key naming, not file encoding. Changing the prefix to end in .parquet does not convert JSON payloads into Parquet; it only changes the object key. The data would still be JSON bytes with a misleading extension, so downstream Parquet readers would fail to parse the files.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.