DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer needs to catalog a growing S3 data lake. New CSV files land in s3://analytics/raw/orders/ with a partition structure year=YYYY/month=MM/day=DD/. The engineer must create an AWS Glue Data Catalog table that automatically recognizes these partitions and requires no crawler runs for future dates. Which approach meets these requirements?
⚠ Common exam trap
The trap here is assuming the Data Catalog must always be refreshed by a crawler or MSCK REPAIR TABLE before new partitions are visible, when partition projection removes that dependency entirely.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create the table with CREATE EXTERNAL TABLE and specify PARTITIONED BY (year string, month string, day string), then enable partition projection in the table properties.
Partition projection makes Athena and Glue compute partition metadata from the table's own properties, so newly arriving year/month/day folders are queryable without a crawler or repair command. The other choices all require periodic human or scheduled action after each new partition lands, which violates the explicit no-crawler-runs requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a Glue crawler with a schedule of every hour to discover new partitions.
Why it's wrong here
A scheduled crawler can discover the new day partitions, but it adds recurring cost and latency, and the question explicitly requires no crawler runs for future dates. It also has to re-scan the S3 path each time, which grows slower as the data lake expands, and there is a window where newly written partitions are not queryable until the next crawler execution completes.
- ✓
Create the table with CREATE EXTERNAL TABLE and specify PARTITIONED BY (year string, month string, day string), then enable partition projection in the table properties.
Why this is correct
Partition projection on an Athena/Glue table lets the engine compute partition locations mathematically from the defined pattern and date range instead of consulting the catalog for each partition. New day, month, or year folders are therefore queryable immediately with no crawler or repair step, which is exactly the automatic, low-maintenance behavior the scenario requires.
- ✗
Create a table using an AWS Glue crawler once, then run MSCK REPAIR TABLE on the table after each new partition is written.
Why it's wrong here
MSCK REPAIR TABLE works for Hive-style partitions but must be run after every new partition write, which means an operational step per day and no true automatic behavior. It can also be slow and expensive on very large tables because it lists and compares all partitions stored in the catalog, so it does not satisfy the requirement of zero ongoing maintenance.
- ✗
Create the table with CREATE EXTERNAL TABLE and add each partition manually with ALTER TABLE ADD PARTITION as files arrive.
Why it's wrong here
Manually adding partitions is a viable one-time technique but requires a human or a script to run ALTER TABLE ADD PARTITION for every new day, so it is the opposite of automatic. It also becomes error-prone at scale and creates a hard dependency on the pipeline that writes the data, which the scenario does not mention and which adds fragile coupling.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.