DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is optimizing an AWS Glue ETL job that reads a large dataset from Amazon S3 and writes to Amazon Redshift. The job currently runs slowly and consumes many DPUs. The engineer wants to improve performance and reduce cost. Which two actions should the engineer take? (Choose two.)
⚠ Common exam trap
The trap here is thinking that adding more DPUs is always the answer to a slow Glue job; the question asks to improve performance and reduce cost, which requires reducing data processed rather than scaling compute.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partition the source data in Amazon S3 and use predicate pushdown in the Glue job
Job bookmarks reduce the data read on subsequent runs by tracking processed files, and partitioning with predicate pushdown reduces the data scanned per run. Together they lower runtime and DPU consumption, improving performance while reducing cost. Increasing DPUs or converting to CSV would increase cost or hurt performance, and disabling auto-scaling removes a cost-saving feature.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of DPUs to the maximum allowed for the job
Why it's wrong here
Adding DPUs can speed up a job, but it also increases cost. The goal is to improve performance and reduce cost, so simply scaling up compute without addressing data volume or partitioning is not the best action. It may also be unnecessary if the job is inefficient due to reading redundant data.
- ✓
Partition the source data in Amazon S3 and use predicate pushdown in the Glue job
Why this is correct
Partitioning the S3 data and using predicate pushdown allows Glue to read only the partitions needed for the query, reducing I/O and the amount of data processed. This improves job performance and lowers DPU usage. It is a targeted optimization that addresses the root cause of slow reads from large datasets.
- ✓
Enable job bookmarks to process only new data on subsequent runs
Why this is correct
Job bookmarks track the state of previously processed data so that subsequent job runs only read new or changed files. This reduces the volume of data read and processed, lowering DPU consumption and runtime. For a recurring ETL job, enabling bookmarks is a standard optimization that directly addresses both performance and cost.
- ✗
Convert the output to CSV instead of Parquet to reduce write time
Why it's wrong here
CSV is row-based and uncompressed by default, which increases storage size and write time compared to columnar Parquet. Converting to CSV would likely worsen performance and increase Redshift load time. Parquet is more efficient for both storage and query, so this action is counterproductive.
- ✗
Disable auto-scaling to keep the job at a fixed capacity
Why it's wrong here
Auto-scaling dynamically adjusts the number of workers based on workload, which can improve performance and reduce cost by scaling down when the job is not fully utilizing resources. Disabling it removes that benefit and can lead to over-provisioning or under-provisioning, so it does not help meet the optimization goals.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.