Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A company runs a data ingestion pipeline that uses AWS Glue to read 500 GB of JSON files from an S3 bucket (s3://raw-data/) every hour. The Glue ETL job transforms the data and writes Parquet files to another S3 bucket (s3://processed-data/). The job is triggered by a time-based CloudWatch Events rule. Recently, the job has started taking over 2 hours to complete, causing delays in downstream processes. The data volume has been consistent, and no changes have been made to the job code or infrastructure. The S3 bucket 's3://raw-data/' receives new files continuously, but the Glue job reads all files in the bucket each run (no incremental processing). The engineer suspects that the job is reprocessing old data. Which action should the engineer take FIRST to reduce the job duration?

⚠ Common exam trap

DEA-C01 often tests the misconception that scaling resources (more DPUs) or optimizing data layout (partitioning) is the first step to improve job performance, when the root cause is reprocessing old data due to missing job bookmarks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable Glue job bookmarking and configure the job to process only new data.

The Glue job is reprocessing all 500 GB of data every hour because it lacks job bookmarking, which tracks previously processed data. Enabling job bookmarking allows Glue to persist state about which files have already been processed, so subsequent runs only read new files. This directly reduces the input volume and thus the job duration without changing the job logic or infrastructure. Since the data volume is consistent and no code changes were made, the issue is purely due to reprocessing old data, making bookmarking the correct first step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable Glue job bookmarking and configure the job to process only new data.

    Why this is correct

    Without job bookmarks, every run rereads the entire s3://raw-data/ prefix, so the 500 GB backlog is reprocessed hourly and runtime balloons. Enabling bookmarks tracks previously processed objects, letting Glue read only newly arrived files, which directly shortens the job.

  • ✗

    Increase the parallelism of the Spark job by repartitioning the data.

    Why it's wrong here

    Repartitioning redistributes records within a single run; it cannot stop the job rereading every historical object under s3://raw-data/ each hour. It is tempting because skewed partitions genuinely slow Spark stages, and repartitioning is the right fix when parallelism is the bottleneck rather than redundant input scanning.

  • ✗

    Add partition pruning by modifying the S3 path to include date-based partitions.

    Why it's wrong here

    Adding date-based partitions only helps if the existing objects already sit under those prefixes; the bucket's current layout is flat, so the path change matches nothing and the job still lists and reads all files. Partition pruning is correct once data is written into a partitioned structure.

  • ✗

    Increase the number of DPUs for the Glue job to 100.

    Why it's wrong here

    Adding DPUs scales compute for the same workload; the job still reads every object under s3://raw-data/ each run, so the reprocessing bottleneck remains. It is tempting because DPU increases suit genuinely compute-bound or skewed jobs, which fits heavy shuffles rather than redundant input reads.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.