DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is configuring an AWS Glue crawler to catalog data stored in Amazon S3. The data is organized as Parquet files under prefixes named by year, month, and day, such as s3://analytics/events/year=2024/month=05/day=17/. Queries in Amazon Athena must use partition pruning to limit scanned data. Which crawler configuration should the engineer choose?
⚠ Common exam trap
Many exam-takers confuse schema-update options such as 'Add new columns only' with partition-update behavior, when partition registration is controlled by the crawler's partition update setting.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a crawler with the S3 path pointing to s3://analytics/events/ and enable 'Update all new and existing partitions with metadata from the table' so partition metadata stays current.
Hive-style prefixes such as year=/month=/day= are recognized as partition keys when an AWS Glue crawler targets the parent S3 prefix. To keep the catalog current as new date prefixes appear, the crawler should be configured to update all new and existing partitions with metadata from the table. Athena then reads the partition keys from the Data Catalog and prunes to only the relevant prefixes, reducing scanned bytes and query cost.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a crawler for each individual day prefix, such as s3://analytics/events/year=2024/month=05/day=17/, so each day becomes its own table.
Why it's wrong here
Creating a separate crawler per day prefix would produce many small tables instead of one partitioned table, which defeats partition pruning and complicates queries. Each table would represent a single day, so a query spanning multiple days would require unions. This approach also creates ongoing maintenance overhead as new days arrive, and it does not leverage Athena's ability to prune partitions within a single table.
- ✓
Create a crawler with the S3 path pointing to s3://analytics/events/ and enable 'Update all new and existing partitions with metadata from the table' so partition metadata stays current.
Why this is correct
The Hive-style year=/month=/day= prefixes are automatically recognized as partitions when the crawler targets the parent prefix. The 'Update all new and existing partitions with metadata from the table' setting ensures that as new date prefixes appear, the crawler adds them and refreshes partition metadata. Athena can then use the partition keys for pruning, scanning only the relevant date ranges instead of the whole bucket.
- ✗
Create a crawler with the S3 path pointing to s3://analytics/events/ and set the 'Schema change policy' to 'Delete tables and columns' so stale partitions are removed.
Why it's wrong here
The 'Delete tables and columns' schema change policy removes tables and columns from the Data Catalog when they are no longer found, which is destructive and unrelated to partition registration. It does not add new partitions for recently created date prefixes, so partition pruning for new data would fail. This setting risks catalog data loss without satisfying the requirement to discover and prune partitions.
- ✗
Create a crawler with the S3 path pointing to s3://analytics/events/ and enable the 'Add new columns only' option so new partitions are detected automatically.
Why it's wrong here
Pointing the crawler at the top-level prefix is correct for discovering the Hive-style partition structure, but 'Add new columns only' is a schema-update policy for table columns, not a partition-discovery setting. It does not control whether new partitions are added. Without the appropriate partition update behavior, newly added date prefixes might not be registered in the Data Catalog, breaking partition pruning for recent data.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.