Courseiva
Data Operations and Support →mediumMultiple Choice

DEA-C01 Data Operations and Support Practice Question

A data engineer has an AWS Glue job that reads JSON files from Amazon S3, applies transformations, and writes Parquet files to another S3 location. The job runs daily and takes about 2 hours. Recently, the job has been failing intermittently with an error indicating that the job bookmark is not being updated correctly, causing duplicate processing of some files. The engineer needs to ensure that only new files are processed on each run. Which action should the engineer take to resolve this issue?

⚠ Common exam trap

The trap here is assuming that performance tuning or output format changes can resolve data duplication, when the root cause is the missing or misconfigured job bookmark feature.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable job bookmarks in the AWS Glue job configuration and ensure that the input source is an S3 data store with the correct path.

AWS Glue job bookmarks maintain state to track data already processed. When enabled with an S3 source, Glue uses the bookmark to process only new files. The intermittent failures and duplicate processing indicate the bookmark is not being updated, likely because it is not enabled or misconfigured. Enabling and correctly configuring job bookmarks resolves the issue.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Modify the AWS Glue job to use a custom transformation that filters files based on their last modified timestamp.

    Why it's wrong here

    A custom timestamp filter is a manual workaround and not the managed bookmark feature. It requires additional code and state management, which is error-prone. Job bookmarks are the native, reliable way to track processed files in AWS Glue, and using a custom filter does not address the underlying bookmark update failure.

  • ✓

    Enable job bookmarks in the AWS Glue job configuration and ensure that the input source is an S3 data store with the correct path.

    Why this is correct

    Job bookmarks are the AWS Glue feature that tracks previously processed data. Enabling them in the job configuration and specifying the correct S3 path allows Glue to persist state and skip already-processed files on subsequent runs. This directly addresses duplicate processing by maintaining a bookmark for the S3 source.

  • ✗

    Change the output format to Parquet and enable compression to reduce the chance of bookmark errors.

    Why it's wrong here

    Output format and compression affect storage and performance, not bookmark tracking. Bookmark errors are related to state management for input sources, not output characteristics. Changing output format does not influence whether Glue correctly identifies new input files, so duplicate processing would persist.

  • ✗

    Increase the number of AWS Glue DPUs to improve job performance and avoid timeouts that cause bookmark failures.

    Why it's wrong here

    Increasing DPUs may speed up the job but does not fix bookmark state issues. The problem is duplicate processing due to bookmark not updating, not performance. More DPUs won't ensure only new files are processed; they just provide more compute resources, which is irrelevant to the bookmark mechanism.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.