DEA-C01 Data Ingestion and Transformation Practice Question
A company runs a data ingestion pipeline that uses AWS Glue to read 500 GB of JSON files from an S3 bucket (s3://raw-data/) every hour. The Glue ETL job transforms the data and writes Parquet files to another S3 bucket (s3://processed-data/). The job is triggered by a time-based CloudWatch Events rule. Recently, the job has started taking over 2 hours to complete, causing delays in downstream processes. The data volume has been consistent, and no changes have been made to the job code or infrastructure. The S3 bucket 's3://raw-data/' receives new files continuously, but the Glue job reads all files in the bucket each run (no incremental processing). The engineer suspects that the job is reprocessing old data. Which action should the engineer take FIRST to reduce the job duration?
⚠ Common exam trap
DEA-C01 often tests the misconception that scaling resources (more DPUs) or optimizing data layout (partitioning) is the first step to improve job performance, when the root cause is reprocessing old data due to missing job bookmarks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Glue job bookmarking and configure the job to process only new data.
The Glue job is reprocessing all 500 GB of data every hour because it lacks job bookmarking, which tracks previously processed data. Enabling job bookmarking allows Glue to persist state about which files have already been processed, so subsequent runs only read new files. This directly reduces the input volume and thus the job duration without changing the job logic or infrastructure. Since the data volume is consistent and no code changes were made, the issue is purely due to reprocessing old data, making bookmarking the correct first step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable Glue job bookmarking and configure the job to process only new data.
Why this is correct
Without job bookmarks, every run rereads the entire s3://raw-data/ prefix, so the 500 GB backlog is reprocessed hourly and runtime balloons. Enabling bookmarks tracks previously processed objects, letting Glue read only newly arrived files, which directly shortens the job.
- ✗
Increase the parallelism of the Spark job by repartitioning the data.
Why it's wrong here
Repartitioning redistributes records within a single run; it cannot stop the job rereading every historical object under s3://raw-data/ each hour. It is tempting because skewed partitions genuinely slow Spark stages, and repartitioning is the right fix when parallelism is the bottleneck rather than redundant input scanning.
- ✗
Add partition pruning by modifying the S3 path to include date-based partitions.
Why it's wrong here
Adding date-based partitions only helps if the existing objects already sit under those prefixes; the bucket's current layout is flat, so the path change matches nothing and the job still lists and reads all files. Partition pruning is correct once data is written into a partitioned structure.
- ✗
Increase the number of DPUs for the Glue job to 100.
Why it's wrong here
Adding DPUs scales compute for the same workload; the job still reads every object under s3://raw-data/ each run, so the reprocessing bottleneck remains. It is tempting because DPU increases suit genuinely compute-bound or skewed jobs, which fits heavy shuffles rather than redundant input reads.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.