DEA-C01 Data Ingestion and Transformation Practice Question
A company uses Amazon Kinesis Data Firehose to ingest data into an S3 bucket. The data is in JSON format and the team wants to convert it to Parquet before storage. Which TWO configurations are required?
⚠ Common exam trap
DEA-C01 often tests whether candidates know that Firehose Parquet conversion is a native two-part configuration (Glue table + Firehose setting) rather than something requiring Lambda, Glue ETL, or Kinesis Data Analytics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a Glue table with the schema of the data.
Option B is correct because Firehose data format conversion relies on the AWS Glue Data Catalog: you must create a Glue table (with the appropriate schema and SerDe) that Firehose references so it knows how to interpret the incoming JSON records. Option E is correct because the actual conversion is enabled in the Firehose delivery stream configuration by turning on data format conversion and setting the output format to Parquet (with the Glue table as the schema source). Option A is not required because Kinesis Data Analytics is for SQL/Flink stream processing, not for Firehose's built-in format conversion. Option C is not required because Firehose performs the JSON-to-Parquet conversion natively via Glue, so a custom Lambda transformation is unnecessary. Option D is not required because Athena is a query service for reading data in S3, not a prerequisite for converting it during ingestion.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Kinesis Data Analytics to transform data to Parquet.
Why it's wrong here
Kinesis Data Analytics performs SQL or Flink stream processing, not Parquet serialisation for Firehose delivery. It is tempting because it sits in the same streaming family, and would be correct when you need windowed aggregations or joins before the data reaches S3.
- ✓
Create a Glue table with the schema of the data.
Why this is correct
Firehose's Parquet conversion relies on the AWS Glue Data Catalog to resolve the source schema, so a Glue table describing the JSON structure is mandatory. Without it, Firehose cannot map incoming records to Parquet columns, and the conversion configuration fails.
- ✗
Configure a Lambda function to convert data on the fly.
Why it's wrong here
Firehose converts JSON to Parquet natively through its record format conversion setting; a Lambda invocation adds latency, cost and operational overhead without being required. It is tempting because Lambda customises records, and would be correct for bespoke transformations Firehose cannot perform itself.
- ✗
Set up an Athena table to read the data.
Why it's wrong here
An Athena table is a query-layer definition over existing S3 objects; it neither converts JSON nor affects Firehose delivery. It is tempting because Athena reads Parquet efficiently, and would be correct for querying converted data after the Firehose configuration is already in place.
- ✓
Enable data format conversion in Firehose and set Output format to Parquet.
Why this is correct
Enabling data format conversion and setting the output format to Parquet activates Firehose's schema-based serialiser, which transforms incoming JSON records into columnar Parquet before delivery. This is the mechanism that satisfies the requirement to store Parquet rather than raw JSON.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.