Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is configuring an AWS Glue crawler to catalog data in an Amazon S3 bucket. The bucket contains CSV files organized in folders by year and month, and new files are added daily. The engineer wants the crawler to detect schema changes automatically and avoid reprocessing unchanged files on subsequent runs. (Choose two.)

⚠ Common exam trap

The trap here is assuming that S3 event notifications or exclude patterns will make the crawler incremental, when only the built-in incremental crawling feature tracks previously crawled paths and files.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable the crawler's incremental crawling feature to identify and process only new folders and files since the last crawl.

Incremental crawling lets the crawler process only new folders and files since the last run, avoiding redundant scanning of unchanged data. The schema change policy, when set to update the table in the Data Catalog, automatically incorporates new columns detected during a crawl. Together these settings satisfy both requirements: detecting schema evolution and skipping already-processed files.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable the crawler's incremental crawling feature to identify and process only new folders and files since the last crawl.

    Why this is correct

    Incremental crawling lets the crawler track previously crawled S3 paths and process only new folders and files on subsequent runs. This directly satisfies the requirement to avoid reprocessing unchanged files, reducing crawl time and cost. Combined with a schema change policy, the crawler keeps the catalog current while minimizing redundant work.

  • ✗

    Set the crawler's S3 target to exclude the folders that have already been cataloged using an exclude pattern.

    Why it's wrong here

    Exclude patterns prevent the crawler from scanning certain paths, but they require manual maintenance as new folders are added. They do not automatically detect schema changes or track which files are new. Using exclude patterns would also risk skipping new data in existing folders, breaking the daily ingestion requirement rather than solving incremental crawling.

  • ✓

    Enable the crawler's schema change policy to update the table definition in the Data Catalog when new columns are detected.

    Why this is correct

    The schema change policy controls what happens when the crawler detects changes such as new columns. Setting it to update the table in the Data Catalog ensures new columns are added to the table definition automatically. This satisfies the requirement to detect schema changes without manual intervention, keeping the catalog aligned with the evolving data in S3.

  • ✗

    Configure the crawler to use S3 event notifications so it runs only when new objects are created.

    Why it's wrong here

    AWS Glue crawlers support scheduled and on-demand runs, and can be triggered by EventBridge rules, but simply configuring S3 event notifications does not make the crawler ignore unchanged files. The crawler still scans the configured S3 path based on its own logic. This option does not directly address avoiding reprocessing of unchanged files or schema change detection.

  • ✗

    Change the crawler's output to create a separate table for each partition folder.

    Why it's wrong here

    Creating a separate table per partition folder fragments the catalog and complicates queries across years and months. AWS Glue crawlers normally create a single table with partition columns when the folder structure follows Hive-style partitioning. This option does not help detect schema changes or skip unchanged files and would make downstream analytics more difficult.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.