DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer needs to transform data in Amazon S3 using AWS Glue. The job must handle schema evolution and partition pruning. Which THREE features should be used?
⚠ Common exam trap
The DEA-C01 exam often tests the distinction between 'incremental processing' (job bookmarks) and 'schema evolution' (Data Catalog + crawlers), leading candidates to incorrectly select job bookmarks for schema changes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AWS Glue Data Catalog
AWS Glue Data Catalog (A) is correct because it stores the table definitions, schemas, and partition metadata that Glue ETL jobs reference, and it is the component that must be updated when schema evolution occurs so the job can read new columns. AWS Glue crawlers (D) are correct because they automatically scan S3 data, detect new or changed columns, and update the Data Catalog with revised schemas and newly discovered partitions, which is exactly how schema evolution is handled. Partition indexes (E) are correct because they accelerate partition pruning by allowing Glue and Athena to look up partitions without listing the entire catalog, dramatically reducing planning time for highly partitioned tables. AWS Glue job bookmarks (B) only track previously processed data to support incremental loads and do not address schema evolution or partition pruning. AWS Glue FindMatches transform (C) is a machine-learning transform for deduplicating and matching records, which is unrelated to schema evolution or partition pruning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
AWS Glue Data Catalog
Why this is correct
The AWS Glue Data Catalog stores table definitions, schemas and partition metadata centrally, so Glue jobs read current schema versions and prune partitions using that metadata. It is the component that lets the job handle evolving schemas and avoid scanning irrelevant partitions.
- ✗
AWS Glue job bookmarks
Why it's wrong here
Job bookmarks track which S3 objects a previous run already processed, preventing reprocessing; they do not reconcile new columns or restrict which partitions a job scans. They are tempting because they are a standard Glue feature, and bookmarks would be correct for incremental loads where only newly arrived files must be transformed.
- ✗
AWS Glue FindMatches transform
Why it's wrong here
FindMatches is a machine learning transform that identifies duplicate or matching records for deduplication; it neither reads evolving schemas nor prunes partitions. It is tempting because Glue transforms are central to this pipeline, and FindMatches would be correct when consolidating customer records that lack a shared unique identifier.
- ✓
AWS Glue crawlers
Why this is correct
Glue crawlers scan S3 data, infer schemas and register new partitions in the Data Catalog, so newly added columns and partitions are detected automatically. This keeps the catalogue current, enabling the job's schema evolution handling and partition pruning.
- ✓
Partition indexes
Why this is correct
Partition indexes accelerate partition pruning by letting AWS Glue query metadata about partition locations without scanning the entire catalog. This satisfies the stem's partition pruning requirement, complementing schema evolution handling and partition projection for efficient S3 data transformation.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.