Courseiva
Data Operations and Support →mediumMultiple Choice

DEA-C01 Data Operations and Support Practice Question

A data engineer manages an AWS Glue ETL job that writes Parquet files to Amazon S3. Downstream Amazon Athena queries started returning duplicate rows after the job was modified to enable job bookmarks. The job reads from an S3 source prefix where new files are appended hourly and the transformation includes a join that reorders records. Which action will most reliably eliminate the duplicate rows while preserving incremental processing?

⚠ Common exam trap

The trap here is assuming that enabling job bookmarks automatically guarantees exactly-once output regardless of how the transformation is written.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Reset the job bookmark state and verify the transformation produces deterministic output for each input file.

Job bookmarks persist state about which S3 objects have been processed. When a transformation changes so that one input object no longer maps cleanly to one output, stale bookmark state causes records to be reprocessed and duplicated. Resetting the bookmark establishes a known baseline, and ensuring the transformation is deterministic per input file keeps each object processed exactly once on subsequent incremental runs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Change the output format from Parquet to CSV so Athena can deduplicate rows during query execution.

    Why it's wrong here

    Athena does not deduplicate rows based on file format. Switching to CSV removes columnar compression benefits, increases scanned bytes, and still returns every row written by Glue. The duplicates originate in the ETL output, so changing the storage format only degrades query performance and cost while leaving the duplicate records fully visible to downstream consumers.

  • ✗

    Disable job bookmarks and reprocess the entire source prefix on every run.

    Why it's wrong here

    Disabling job bookmarks forces the job to read every object in the source prefix on each run. That guarantees the join output is rewritten in full, which does remove duplicates in the target, but it abandons incremental processing entirely and increases runtime and cost linearly with data volume. It does not address the root cause, which is that the bookmark state no longer matches the transformation semantics.

  • ✓

    Reset the job bookmark state and verify the transformation produces deterministic output for each input file.

    Why this is correct

    Job bookmarks track which S3 objects were already processed, so when a transformation changes the input-to-output relationship, stale bookmark state can cause overlaps. Resetting the bookmark forces a clean reprocessing baseline, and making the transformation deterministic per input file ensures each source object maps to exactly one output. Together these restore correct incremental behavior without reprocessing everything on every run.

  • ✗

    Increase the number of Glue DPUs allocated to the job so the join completes in a single task.

    Why it's wrong here

    DPU count affects parallelism and memory available to the Spark executors, not the logical mapping between input objects and output rows. Adding workers will not reconcile bookmark state or prevent the same source record from being emitted twice. The duplicate rows stem from processing semantics, so scaling compute capacity leaves the duplication intact while only making the job run faster.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.