Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is building an AWS Glue ETL job that reads records from an Amazon Kinesis Data Stream and writes them to Amazon S3 in Parquet format. The job must checkpoint its progress so that it can resume without reprocessing data after a failure. Which AWS Glue mechanism should the engineer configure to track the stream position?

⚠ Common exam trap

The trap here is assuming that Kinesis stream retention or retry triggers provide checkpointing, when only Glue job bookmarks persist the last processed position.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable job bookmarks with the transformation_ctx parameter set on the create_data_frame_from_options call.

Glue job bookmarks maintain state between runs, storing the last processed Kinesis sequence number per shard. With transformation_ctx set on the read operation, the job can resume from where it stopped after a failure, avoiding duplicate writes to S3. The other options address throughput, concurrency, or retry orchestration, none of which track record-level progress.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use AWS Glue Studio's auto-generated script and enable the 'retry on failure' option in the job's trigger.

    Why it's wrong here

    A retry trigger re-runs the job after a failure but does not know which records were already processed. Without bookmarks, the retried run re-reads the full stream window, producing duplicates. Retry logic addresses fault tolerance at the orchestration level, not data-level checkpointing, so it is insufficient here.

  • ✓

    Enable job bookmarks with the transformation_ctx parameter set on the create_data_frame_from_options call.

    Why this is correct

    Job bookmarks persist the state of the last processed record per source. When reading from Kinesis, Glue stores the stream sequence number and shard information, so a re-run resumes from the checkpoint. Setting transformation_ctx uniquely identifies the source in the bookmark store, preventing conflicting state when a job reads multiple streams.

  • ✗

    Set the Glue job's MaxConcurrentRuns parameter to 1 and let the job restart from the beginning on failure.

    Why it's wrong here

    MaxConcurrentRuns limits how many instances of the job can run simultaneously; it does not track which records were processed. Restarting from the beginning would duplicate data in S3 and waste compute. This parameter controls concurrency, not ingestion progress, so it cannot satisfy the no-reprocessing requirement.

  • ✗

    Configure a Kinesis Data Stream consumer with enhanced fan-out and rely on the stream's 24-hour retention to replay records.

    Why it's wrong here

    Enhanced fan-out increases the read throughput per consumer and reduces latency, but it does not persist ETL progress. Replaying from the stream's retention window would reprocess records already written to S3 and would fail once the retention period expires. It does not provide the resumable checkpoint behavior the scenario requires.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.