Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A healthcare company is ingesting patient data from a legacy system into an Amazon S3 data lake using AWS Glue. The legacy system produces CSV files with inconsistent schemas (columns may appear or disappear in different files). The data engineer needs to create a Glue ETL job that can handle schema evolution and transform the data into a standardized parquet format. The job should also be able to process new files as they arrive. Which approach should the data engineer use?

⚠ Common exam trap

The trap is assuming a static schema or standard Spark DataFrame can handle schema evolution; candidates must recognize DynamicFrames as the Glue-native solution for inconsistent schemas.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use AWS Glue DynamicFrames to read the CSV files and apply transformations using resolveChoice and applyMapping.

AWS Glue DynamicFrames are designed to handle schema evolution and inconsistent data. Using resolveChoice to handle columns that appear/disappear and applyMapping to standardize the schema allows the job to process files with varying schemas and output consistent Parquet.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use AWS Glue crawlers to create a schema in the Data Catalog and then use a standard Spark DataFrame for transformation.

    Why it's wrong here

    Crawlers infer one schema per table, so files with differing columns either overwrite the catalog definition or land in separate tables; a standard Spark DataFrame then reads only the matching schema. Crawlers suit stable sources, but evolving CSV schemas require Glue dynamic frames with mergeSchema and job bookmarks.

  • ✓

    Use AWS Glue DynamicFrames to read the CSV files and apply transformations using resolveChoice and applyMapping.

    Why this is correct

    DynamicFrames natively accommodate schema evolution: resolveChoice reconciles columns that appear or disappear across files by casting or dropping them, while applyMapping standardises surviving fields before writing Parquet. This satisfies the inconsistent-schema constraint, and Glue job bookmarks let the same job process newly arrived files incrementally.

  • ✗

    Use a Python shell job in Glue to manually parse each file and write to parquet.

    Why it's wrong here

    A Python shell job runs single-node Python without the Spark DataFrame reader's mergeSchema and schema-evolution handling, so inconsistent CSV columns cannot be reconciled into Parquet. Python shell suits lightweight scripting, such as small file conversions, but not distributed ETL with evolving schemas and continuous new-file processing.

  • ✗

    Use a Glue ETL job with a static schema defined in the script and ignore files that don't match.

    Why it's wrong here

    A static schema in the script cannot accommodate columns appearing or disappearing; files lacking declared columns fail or are dropped, losing patient records. Static schemas suit stable, well-governed sources, but schema evolution here demands Glue's dynamic frame with mergeSchema and Data Catalog updates.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.