Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer maintains an AWS Glue job that incrementally processes new files in Amazon S3 using job bookmarks. After a schema change in the source data added a new column, the engineer updated the Glue Data Catalog table. Subsequent job runs still process only previously seen files and ignore newly arrived objects. The engineer verifies that new files exist in the prefix and that the bookmark state was not reset. Which factor most likely explains why new files are being skipped?

⚠ Common exam trap

The trap here is assuming bookmarks only look at file timestamps, when they actually key on transformation context that script or catalog edits can invalidate.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Job bookmarks use the source path and a transformation context key; changing the job script or the catalog table can invalidate the bookmark state.

Job bookmarks persist state keyed to the source path, transformation context, and job parameters. Modifying the job script or the referenced Data Catalog table alters that context, which can cause the bookmark to skip newly arrived objects despite their presence. Resetting the bookmark or recreating the job restores expected incremental behavior. Format, IAM, and S3 consistency are not the cause here.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon S3 eventual consistency delays object visibility, so bookmarks never see files written within the same hour.

    Why it's wrong here

    Amazon S3 provides strong read-after-write consistency for new object PUTs, so newly written files are immediately visible. Eventual consistency is not the cause of files being skipped across multiple job runs. This explanation misstates S3 consistency guarantees and does not account for the schema-change context.

  • ✗

    Job bookmarks only support Amazon S3 sources when the files are in Parquet format, so JSON files are silently skipped.

    Why it's wrong here

    Job bookmarks support multiple S3 formats including JSON, CSV, and Parquet. The format itself does not disable bookmark tracking. The described symptom of skipping new files after a schema change points to bookmark state invalidation, not a format limitation, so this explanation does not fit the scenario.

  • ✗

    The Glue job's IAM role lacks s3:ListBucket permission, so the job cannot detect newly arrived objects.

    Why it's wrong here

    A missing s3:ListBucket permission would cause an access denied failure rather than a silent skip of new files. The scenario states the engineer verified new files exist, implying listing works. Bookmark behavior, not IAM, governs which objects are considered unprocessed.

  • ✓

    Job bookmarks use the source path and a transformation context key; changing the job script or the catalog table can invalidate the bookmark state.

    Why this is correct

    Glue job bookmarks track processed data using a key derived from the source path, transformation context, and related job parameters. Editing the script or altering the catalog table can change that context, causing the bookmark to behave unexpectedly and skip new files. Resetting or recreating the bookmark after such changes restores correct incremental processing.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.