Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer runs an AWS Glue job that reads from an Amazon Kinesis Data Stream and writes to Amazon S3. The job must process records in order per shard and must checkpoint progress so it can resume after a failure without reprocessing all records. Which Glue configuration BEST supports this?

⚠ Common exam trap

Many exam-takers confuse `--starting_position` on each run with true checkpointing, when resumption without reprocessing depends on job bookmarks being enabled.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Glue streaming ETL job type with `--starting_position` set to `TRIM_HORIZON` and rely on Glue's built-in checkpointing in the job bookmarks.

Glue streaming ETL jobs natively support Kinesis Data Streams and use job bookmarks to track progress per shard. Setting `--starting_position` to `TRIM_HORIZON` establishes the initial read point, and bookmarks checkpoint processed records so a failed job resumes without reprocessing everything. Batch jobs, disabled bookmarks, and custom checkpoint tables either lose data, reprocess, or add unnecessary overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a Glue streaming ETL job with a custom checkpoint table in Amazon DynamoDB and set `--starting_position` to `TRIM_HORIZON`.

    Why it's wrong here

    Glue streaming jobs already provide checkpointing through job bookmarks, so a custom DynamoDB checkpoint table adds unnecessary complexity and operational overhead. While custom checkpointing is possible in custom Spark Streaming applications, it is not the best choice for a Glue job. The requirement is met by the built-in mechanism, making this approach redundant and harder to maintain.

  • ✗

    Use a Glue streaming ETL job with `--starting_position` set to `LATEST` and disable job bookmarks to always read new records.

    Why it's wrong here

    Disabling job bookmarks removes the checkpointing that allows resumption after failure. Without bookmarks, a restart could reprocess records or skip them depending on the starting position. Setting `LATEST` on restart would skip records produced during downtime, violating the requirement to resume without reprocessing all records while maintaining order.

  • ✓

    Use the Glue streaming ETL job type with `--starting_position` set to `TRIM_HORIZON` and rely on Glue's built-in checkpointing in the job bookmarks.

    Why this is correct

    Glue streaming ETL jobs support Kinesis Data Streams as a source and use job bookmarks to track processed records. Setting `--starting_position` to `TRIM_HORIZON` defines where to start on the first run, and bookmarks checkpoint progress so subsequent runs resume from the last processed record. This provides ordered, per-shard processing and fault-tolerant resumption without custom checkpoint code.

  • ✗

    Use a standard Glue batch job with the Kinesis connector and set `--starting_position` to `LATEST` for each run.

    Why it's wrong here

    A standard batch job does not maintain streaming checkpoints and would not continuously process the stream. Setting `--starting_position` to `LATEST` on each run would skip records produced between runs, causing data loss. Batch jobs also do not provide per-shard ordering guarantees for a stream source in the way streaming ETL jobs do.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.