DEA-C01 Data Ingestion and Transformation Practice Question
A company is using AWS Glue to catalog data in Amazon S3. The data is stored in CSV format, but the schema is not consistent across all files. Which TWO actions can the company take to handle schema evolution and ensure the Glue Data Catalog is up to date? (Choose TWO.)
⚠ Common exam trap
DEA-C01 often tests the features of Glue crawlers for schema evolution, and candidates may overlook the need for both enabling schema update and scheduling the crawler.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the Glue crawler to update the table schema on each run.
Option A is correct because configuring the Glue crawler to update the table schema on each run allows it to detect new columns or changed data types in the CSV files and automatically revise the Data Catalog table definition, which is essential when schemas vary across files. Option D is correct because scheduling the crawler to run periodically ensures that any schema changes introduced by new or modified files are detected and reflected in the Glue Data Catalog in a timely, automated manner. Together, these two actions provide automated, recurring schema evolution handling. Option B is not ideal because manual updates are error-prone and do not scale, and the question asks for actions the company can take to handle schema evolution automatically. Option C is incorrect because disabling schema update prevents the crawler from adapting to schema changes, and manual partition addition does not address evolving column structures. Option E is incorrect because enforcing a single fixed schema contradicts the scenario where schemas are already inconsistent and does not help update the catalog for existing variation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure the Glue crawler to update the table schema on each run.
Why this is correct
Configuring the crawler to update the table schema on each run lets it revise column definitions in the Glue Data Catalog when it encounters differing CSV structures, directly handling schema evolution rather than leaving stale definitions in place.
- ✗
Manually update the Glue Data Catalog tables whenever the schema changes.
Why it's wrong here
Manual catalog edits do not scale and drift out of sync as new files arrive, so the catalog quickly becomes stale. It is tempting because hand-editing gives precise control over column definitions, which suits small, stable datasets where crawler inference misreads types and a human correction is genuinely warranted.
- ✗
Disable schema update in the crawler and add partitions manually.
Why it's wrong here
Disabling schema update stops the crawler detecting new columns, and manual partition registration cannot reconcile inconsistent CSV headers, leaving the catalog incomplete. It is tempting because disabling updates prevents crawler-induced schema churn, which is correct when a table's structure is genuinely fixed and partitions are managed externally.
- ✓
Schedule the Glue crawler to run periodically to detect changes.
Why this is correct
Scheduling the crawler periodically re-scans S3, detecting newly arrived files and changed schemas so the Glue Data Catalog reflects current structure. This satisfies the requirement to keep the catalog up to date as CSV files with inconsistent schemas continue to land.
- ✗
Require all data producers to use a single fixed schema.
Why it's wrong here
Enforcing one fixed schema across producers is an organisational mandate, not a Glue mechanism, and cannot accommodate files already written with divergent headers. It is tempting because standardising upstream eliminates schema drift at source, which is the right long-term fix when producers can actually be governed.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.