Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A company is using AWS Glue to catalog data in Amazon S3. The data is stored in CSV format, but the schema is not consistent across all files. Which TWO actions can the company take to handle schema evolution and ensure the Glue Data Catalog is up to date? (Choose TWO.)

⚠ Common exam trap

DEA-C01 often tests the features of Glue crawlers for schema evolution, and candidates may overlook the need for both enabling schema update and scheduling the crawler.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the Glue crawler to update the table schema on each run.

Option A is correct because configuring the Glue crawler to update the table schema on each run allows it to detect new columns or changed data types in the CSV files and automatically revise the Data Catalog table definition, which is essential when schemas vary across files. Option D is correct because scheduling the crawler to run periodically ensures that any schema changes introduced by new or modified files are detected and reflected in the Glue Data Catalog in a timely, automated manner. Together, these two actions provide automated, recurring schema evolution handling. Option B is not ideal because manual updates are error-prone and do not scale, and the question asks for actions the company can take to handle schema evolution automatically. Option C is incorrect because disabling schema update prevents the crawler from adapting to schema changes, and manual partition addition does not address evolving column structures. Option E is incorrect because enforcing a single fixed schema contradicts the scenario where schemas are already inconsistent and does not help update the catalog for existing variation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure the Glue crawler to update the table schema on each run.

    Why this is correct

    Configuring the crawler to update the table schema on each run lets it revise column definitions in the Glue Data Catalog when it encounters differing CSV structures, directly handling schema evolution rather than leaving stale definitions in place.

  • ✗

    Manually update the Glue Data Catalog tables whenever the schema changes.

    Why it's wrong here

    Manual catalog edits do not scale and drift out of sync as new files arrive, so the catalog quickly becomes stale. It is tempting because hand-editing gives precise control over column definitions, which suits small, stable datasets where crawler inference misreads types and a human correction is genuinely warranted.

  • ✗

    Disable schema update in the crawler and add partitions manually.

    Why it's wrong here

    Disabling schema update stops the crawler detecting new columns, and manual partition registration cannot reconcile inconsistent CSV headers, leaving the catalog incomplete. It is tempting because disabling updates prevents crawler-induced schema churn, which is correct when a table's structure is genuinely fixed and partitions are managed externally.

  • ✓

    Schedule the Glue crawler to run periodically to detect changes.

    Why this is correct

    Scheduling the crawler periodically re-scans S3, detecting newly arrived files and changed schemas so the Glue Data Catalog reflects current structure. This satisfies the requirement to keep the catalog up to date as CSV files with inconsistent schemas continue to land.

  • ✗

    Require all data producers to use a single fixed schema.

    Why it's wrong here

    Enforcing one fixed schema across producers is an organisational mandate, not a Glue mechanism, and cannot accommodate files already written with divergent headers. It is tempting because standardising upstream eliminates schema drift at source, which is the right long-term fix when producers can actually be governed.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.