Courseiva

DEA-C01 Data Operations and Support Practice Question

A data engineer is using AWS Glue to process data from an Amazon Kinesis Data Stream. The Glue job is configured to run every 15 minutes and uses job bookmarks to track processed data. Recently, the job started reprocessing old data, leading to duplicate records in the target Amazon S3 bucket. The engineer verifies that the job bookmark is enabled and the job is not being run manually. Which TWO actions should the engineer take to resolve the duplicate processing issue? (Choose two.)

⚠ Common exam trap

The trap here is assuming that enabling job bookmarks is sufficient to prevent duplicates, when in fact bookmark key changes or concurrent runs can still cause reprocessing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check that the Glue job's bookmark key is not being changed between runs, as a changed bookmark key resets the bookmark state.

Duplicate processing with AWS Glue job bookmarks often occurs when the bookmark key changes or when concurrent runs interfere with bookmark updates. Verifying the bookmark key remains constant and preventing concurrent runs with the same bookmark are key steps. These actions ensure the bookmark state is preserved and correctly updated, preventing reprocessing of already-processed data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Check that the Glue job's bookmark key is not being changed between runs, as a changed bookmark key resets the bookmark state.

    Why this is correct

    The job bookmark key is used to identify the bookmark state. If the key changes, the job treats it as a new job and starts processing from the beginning, causing duplicates. Ensuring the bookmark key remains consistent is crucial for proper bookmark functionality.

  • ✗

    Increase the Kinesis stream's retention period to 168 hours to allow the Glue job more time to process data.

    Why it's wrong here

    Increasing the retention period allows data to remain in the stream longer, but it does not address the duplicate processing issue. The problem is that the job is reprocessing data it has already processed, which is a bookmark management issue, not a data retention issue.

  • ✗

    Configure the Glue job to use the '--job-bookmark-option' parameter set to 'job-bookmark-enable'.

    Why it's wrong here

    The '--job-bookmark-option' parameter is used to enable or disable job bookmarks. Setting it to 'job-bookmark-enable' is the default when bookmarks are enabled, but it does not fix the issue if the bookmark is already enabled. The problem is likely due to bookmark key changes or concurrent runs, not the enablement itself.

  • ✓

    Ensure that the Glue job is not being run concurrently with the same job bookmark, as concurrent runs can cause duplicate processing.

    Why this is correct

    Concurrent runs of the same Glue job with the same bookmark can interfere with each other's bookmark updates, leading to reprocessing of data. AWS Glue does not support concurrent runs with job bookmarks; each run should complete before the next starts to maintain bookmark integrity.

  • ✗

    Verify that the Kinesis stream's shard iterator type is set to LATEST in the Glue job's connection options.

    Why it's wrong here

    Setting the shard iterator type to LATEST would cause the job to start reading from the latest records, skipping older ones, but it would not prevent reprocessing of data already processed. Job bookmarks are designed to track progress; changing the iterator type could cause data loss or duplication depending on the bookmark state.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.