DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes 500 GB of data. The engineer notices that the job takes several hours and wants to optimize performance. The data is stored in Parquet format and partitioned by date. Which optimization should the engineer implement to improve the job's performance?
⚠ Common exam trap
The trap here is assuming that simply adding more DPUs will solve performance problems, when actually data layout optimizations like predicate pushdown and partition pruning often yield greater benefits for partitioned Parquet data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use predicate pushdown to filter data at the source and partition pruning to read only necessary partitions.
Predicate pushdown and partition pruning are key optimizations for AWS Glue jobs reading partitioned Parquet data. They minimize the amount of data scanned by pushing filters to the source and skipping irrelevant partitions. This reduces I/O and compute time, directly addressing the performance issue without unnecessary cost increases.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the Parquet data to CSV to improve read performance.
Why it's wrong here
Parquet is a columnar format optimized for analytics, offering better compression and faster reads than CSV. Converting to CSV would degrade performance and increase storage size. CSV is row-based and does not support predicate pushdown as efficiently. This change would likely slow down the job and increase costs.
- ✓
Use predicate pushdown to filter data at the source and partition pruning to read only necessary partitions.
Why this is correct
Predicate pushdown and partition pruning allow Glue to read only the relevant partitions and rows from S3, reducing I/O and the amount of data processed. Since the data is partitioned by date, the job can target specific partitions. This significantly speeds up the job and lowers cost. It is a best practice for large datasets in Parquet.
- ✗
Increase the number of DPUs for the Glue job to scale horizontally.
Why it's wrong here
Increasing DPUs can improve performance by adding more compute resources, but it is not the most effective optimization for a job reading partitioned Parquet data. Without addressing data layout and pushdown, simply adding DPUs may not yield linear speedup and increases cost. The job may still be I/O bound due to inefficient reads.
- ✗
Enable job bookmarks to track processed data and avoid reprocessing.
Why it's wrong here
Job bookmarks help avoid reprocessing old data, which is useful for incremental loads, but the scenario describes a daily job that processes 500 GB, likely processing new partitions. Bookmarks do not optimize the transformation of that day's data. They reduce redundant work across runs, not the performance of a single run. The job still needs to read and transform all new data.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.