DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is configuring an AWS Glue crawler to catalog data in an Amazon S3 bucket. The bucket contains CSV files organized in folders by year and month, and new files are added daily. The engineer wants the crawler to detect schema changes automatically and avoid reprocessing unchanged files on subsequent runs. (Choose two.)
⚠ Common exam trap
The trap here is assuming that S3 event notifications or exclude patterns will make the crawler incremental, when only the built-in incremental crawling feature tracks previously crawled paths and files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable the crawler's incremental crawling feature to identify and process only new folders and files since the last crawl.
Incremental crawling lets the crawler process only new folders and files since the last run, avoiding redundant scanning of unchanged data. The schema change policy, when set to update the table in the Data Catalog, automatically incorporates new columns detected during a crawl. Together these settings satisfy both requirements: detecting schema evolution and skipping already-processed files.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable the crawler's incremental crawling feature to identify and process only new folders and files since the last crawl.
Why this is correct
Incremental crawling lets the crawler track previously crawled S3 paths and process only new folders and files on subsequent runs. This directly satisfies the requirement to avoid reprocessing unchanged files, reducing crawl time and cost. Combined with a schema change policy, the crawler keeps the catalog current while minimizing redundant work.
- ✗
Set the crawler's S3 target to exclude the folders that have already been cataloged using an exclude pattern.
Why it's wrong here
Exclude patterns prevent the crawler from scanning certain paths, but they require manual maintenance as new folders are added. They do not automatically detect schema changes or track which files are new. Using exclude patterns would also risk skipping new data in existing folders, breaking the daily ingestion requirement rather than solving incremental crawling.
- ✓
Enable the crawler's schema change policy to update the table definition in the Data Catalog when new columns are detected.
Why this is correct
The schema change policy controls what happens when the crawler detects changes such as new columns. Setting it to update the table in the Data Catalog ensures new columns are added to the table definition automatically. This satisfies the requirement to detect schema changes without manual intervention, keeping the catalog aligned with the evolving data in S3.
- ✗
Configure the crawler to use S3 event notifications so it runs only when new objects are created.
Why it's wrong here
AWS Glue crawlers support scheduled and on-demand runs, and can be triggered by EventBridge rules, but simply configuring S3 event notifications does not make the crawler ignore unchanged files. The crawler still scans the configured S3 path based on its own logic. This option does not directly address avoiding reprocessing of unchanged files or schema change detection.
- ✗
Change the crawler's output to create a separate table for each partition folder.
Why it's wrong here
Creating a separate table per partition folder fragments the catalog and complicates queries across years and months. AWS Glue crawlers normally create a single table with partition columns when the folder structure follows Hive-style partitioning. This option does not help detect schema changes or skip unchanged files and would make downstream analytics more difficult.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.