DEA-C01 Data Ingestion and Transformation Practice Question
A healthcare company is ingesting patient data from a legacy system into an Amazon S3 data lake using AWS Glue. The legacy system produces CSV files with inconsistent schemas (columns may appear or disappear in different files). The data engineer needs to create a Glue ETL job that can handle schema evolution and transform the data into a standardized parquet format. The job should also be able to process new files as they arrive. Which approach should the data engineer use?
⚠ Common exam trap
The trap is assuming a static schema or standard Spark DataFrame can handle schema evolution; candidates must recognize DynamicFrames as the Glue-native solution for inconsistent schemas.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS Glue DynamicFrames to read the CSV files and apply transformations using resolveChoice and applyMapping.
AWS Glue DynamicFrames are designed to handle schema evolution and inconsistent data. Using resolveChoice to handle columns that appear/disappear and applyMapping to standardize the schema allows the job to process files with varying schemas and output consistent Parquet.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use AWS Glue crawlers to create a schema in the Data Catalog and then use a standard Spark DataFrame for transformation.
Why it's wrong here
Crawlers infer one schema per table, so files with differing columns either overwrite the catalog definition or land in separate tables; a standard Spark DataFrame then reads only the matching schema. Crawlers suit stable sources, but evolving CSV schemas require Glue dynamic frames with mergeSchema and job bookmarks.
- ✓
Use AWS Glue DynamicFrames to read the CSV files and apply transformations using resolveChoice and applyMapping.
Why this is correct
DynamicFrames natively accommodate schema evolution: resolveChoice reconciles columns that appear or disappear across files by casting or dropping them, while applyMapping standardises surviving fields before writing Parquet. This satisfies the inconsistent-schema constraint, and Glue job bookmarks let the same job process newly arrived files incrementally.
- ✗
Use a Python shell job in Glue to manually parse each file and write to parquet.
Why it's wrong here
A Python shell job runs single-node Python without the Spark DataFrame reader's mergeSchema and schema-evolution handling, so inconsistent CSV columns cannot be reconciled into Parquet. Python shell suits lightweight scripting, such as small file conversions, but not distributed ETL with evolving schemas and continuous new-file processing.
- ✗
Use a Glue ETL job with a static schema defined in the script and ignore files that don't match.
Why it's wrong here
A static schema in the script cannot accommodate columns appearing or disappearing; files lacking declared columns fail or are dropped, losing patient records. Static schemas suit stable, well-governed sources, but schema evolution here demands Glue's dynamic frame with mergeSchema and Data Catalog updates.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.