Courseiva

DEA-C01 Data Operations and Support Practice Question

A data engineer is responsible for a data pipeline that uses Amazon S3 as a data lake, AWS Glue for ETL, and Amazon Athena for ad-hoc queries. The pipeline ingests CSV files from an external partner via SFTP into an S3 bucket. The files are then processed by a Glue job that converts them to Parquet and writes to a separate S3 bucket partitioned by date. The Glue job runs daily and is triggered by a scheduled CloudWatch Events rule. Recently, the data engineer noticed that some days the Glue job fails because of memory errors, and on those days the Athena queries that rely on the data return incomplete results. The engineer needs to ensure that the pipeline is resilient and that Athena queries always see a complete view of the data, even if the Glue job fails mid-run. The engineer also needs to minimize re-processing of data. Which course of action should the engineer take?

⚠ Common exam trap

DEA-C01 often tests whether candidates conflate 'fix the failure' (more workers, retries) with 'ensure atomic visibility' — the trap is choosing a scaling fix that doesn't address partial-data exposure to Athena.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Modify the Glue job to use job bookmarks for incremental processing and write the Parquet output to a temporary location, then use an S3 copy operation to move the data into the final partitioned location only after the job completes successfully.

Writing Parquet output to a temporary location and only copying it into the final partitioned path after the job succeeds ensures Athena never sees partial data from a failed run. Job bookmarks enable incremental processing so already-processed files aren't reprocessed, minimizing re-work. This combination directly addresses both resilience (atomic visibility) and efficiency (no duplicate processing).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of workers and the worker type to G.2X to handle the memory errors, and enable job retries.

    Why it's wrong here

    Adding workers and retries only reduces the chance of failure; a mid-run failure still leaves partial Parquet output that Athena reads, so queries remain incomplete. It is tempting because memory errors are the stated symptom, but scaling does not provide atomic commit of the job's output.

  • ✗

    Replace the Glue job with an AWS Lambda function that processes the CSV files and writes Parquet to S3, and use S3 Event Notifications to trigger the function.

    Why it's wrong here

    Lambda has a 15-minute execution limit and no native bookmarking, so large CSV conversions would time out and reprocess files, worsening the incomplete-data problem. It is tempting because S3 event triggers give immediate processing, but Lambda suits small, short transformations rather than daily bulk ETL.

  • ✓

    Modify the Glue job to use job bookmarks for incremental processing and write the Parquet output to a temporary location, then use an S3 copy operation to move the data into the final partitioned location only after the job completes successfully.

    Why this is correct

    Writing Parquet to a temporary prefix and copying only after successful completion keeps the final partitioned location free of partial output, so Athena never reads incomplete data. Glue job bookmarks track previously processed S3 objects, satisfying the minimise re-processing constraint by skipping already-ingested files on retry.

  • ✗

    Use Athena partition projection to automatically discover partitions and set up a retry mechanism using AWS Step Functions.

    Why it's wrong here

    Partition projection only speeds partition discovery in Athena; it does not make partially written Parquet data visible atomically, and Step Functions retries re-run the whole job rather than resuming. It is tempting because projection reduces query latency and Step Functions orchestrates workflows, but neither addresses incomplete writes.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.