DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is configuring an AWS Glue crawler to catalog CSV files stored in Amazon S3. The files are organized in prefixes by year and month, and the engineer wants the crawler to detect new partitions automatically and avoid re-crawling unchanged partitions. (Choose two.)
⚠ Common exam trap
The trap here is believing that schema change policies or partition indexes control crawler partition discovery, when incremental crawling and Hive-style layout do.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the crawler to use incremental crawling so it processes only folders added since the last crawl.
Incremental crawling limits each crawler run to newly added folders, so previously cataloged partitions are not rescanned, reducing time and cost. Hive-style key=value prefixes make the year and month directories recognizable as partitions so the crawler can populate them correctly in the Data Catalog. Together they deliver automatic partition detection without reprocessing unchanged data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the crawler's schema change policy to update the table definition in the Data Catalog.
Why it's wrong here
The schema change policy controls how the crawler handles changes to the table schema, such as adding or deleting columns. It does not govern partition detection or incremental crawling of new prefixes. While useful for keeping the schema current, it does not address the requirement to detect new partitions efficiently or to skip unchanged partitions during subsequent crawls.
- ✓
Configure the crawler to use incremental crawling so it processes only folders added since the last crawl.
Why this is correct
Incremental crawling lets a Glue crawler examine only new partitions or folders that appeared since the previous run, rather than rescanning the entire S3 prefix. This directly reduces crawl time and cost for a partitioned layout organized by year and month, satisfying the requirement to detect new partitions automatically while avoiding re-crawling unchanged partitions.
- ✓
Ensure the S3 prefixes follow a Hive-style partition naming convention with key=value pairs.
Why this is correct
Hive-style partitioning such as year=2024/month=03 allows the Glue crawler to recognize the directory structure as partitions and populate partition columns in the Data Catalog table. Combined with incremental crawling, this enables automatic detection of new partitions by year and month, which is exactly the layout the scenario describes and the behavior the engineer needs.
- ✗
Add a path in the crawler configuration that points to each year and month prefix individually.
Why it's wrong here
Listing every year and month prefix separately is a manual, brittle approach that requires updates as new periods arrive. It does not automatically detect new partitions and would still cause the crawler to examine specified paths regardless of whether data changed. This increases maintenance overhead and contradicts the goal of automatic partition detection.
- ✗
Create a partition index on the Data Catalog table to speed up partition filtering.
Why it's wrong here
A partition index improves query performance when filtering partitions in services like Athena, but it does not change how the crawler discovers or re-crawls partitions. The scenario concerns crawler behavior rather than query acceleration. Adding a partition index would not prevent the crawler from rescanning unchanged prefixes, so it does not meet either stated requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.