Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer needs to transform data in Amazon S3 using AWS Glue. The job must handle schema evolution and partition pruning. Which THREE features should be used?

⚠ Common exam trap

The DEA-C01 exam often tests the distinction between 'incremental processing' (job bookmarks) and 'schema evolution' (Data Catalog + crawlers), leading candidates to incorrectly select job bookmarks for schema changes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

AWS Glue Data Catalog

AWS Glue Data Catalog (A) is correct because it stores the table definitions, schemas, and partition metadata that Glue ETL jobs reference, and it is the component that must be updated when schema evolution occurs so the job can read new columns. AWS Glue crawlers (D) are correct because they automatically scan S3 data, detect new or changed columns, and update the Data Catalog with revised schemas and newly discovered partitions, which is exactly how schema evolution is handled. Partition indexes (E) are correct because they accelerate partition pruning by allowing Glue and Athena to look up partitions without listing the entire catalog, dramatically reducing planning time for highly partitioned tables. AWS Glue job bookmarks (B) only track previously processed data to support incremental loads and do not address schema evolution or partition pruning. AWS Glue FindMatches transform (C) is a machine-learning transform for deduplicating and matching records, which is unrelated to schema evolution or partition pruning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    AWS Glue Data Catalog

    Why this is correct

    The AWS Glue Data Catalog stores table definitions, schemas and partition metadata centrally, so Glue jobs read current schema versions and prune partitions using that metadata. It is the component that lets the job handle evolving schemas and avoid scanning irrelevant partitions.

  • ✗

    AWS Glue job bookmarks

    Why it's wrong here

    Job bookmarks track which S3 objects a previous run already processed, preventing reprocessing; they do not reconcile new columns or restrict which partitions a job scans. They are tempting because they are a standard Glue feature, and bookmarks would be correct for incremental loads where only newly arrived files must be transformed.

  • ✗

    AWS Glue FindMatches transform

    Why it's wrong here

    FindMatches is a machine learning transform that identifies duplicate or matching records for deduplication; it neither reads evolving schemas nor prunes partitions. It is tempting because Glue transforms are central to this pipeline, and FindMatches would be correct when consolidating customer records that lack a shared unique identifier.

  • ✓

    AWS Glue crawlers

    Why this is correct

    Glue crawlers scan S3 data, infer schemas and register new partitions in the Data Catalog, so newly added columns and partitions are detected automatically. This keeps the catalogue current, enabling the job's schema evolution handling and partition pruning.

  • ✓

    Partition indexes

    Why this is correct

    Partition indexes accelerate partition pruning by letting AWS Glue query metadata about partition locations without scanning the entire catalog. This satisfies the stem's partition pruning requirement, complementing schema evolution handling and partition projection for efficient S3 data transformation.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.