DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year, month, and day. The engineer notices that the Glue job is taking a long time and consuming many DPUs. The job reads all partitions, filters the data, and writes the result to another S3 location. The engineer wants to optimize the job to process only the required partitions and reduce cost. Which action should the engineer take?
⚠ Common exam trap
The trap here is thinking that job bookmarks or more DPUs will reduce the data read; bookmarks only avoid reprocessing, and more DPUs increase cost without reducing data volume.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use predicate pushdown in the Glue job by specifying a push_down_predicate when reading from the AWS Glue Data Catalog.
Predicate pushdown in AWS Glue allows the job to read only the partitions that match a specified condition, reducing data scanned and DPU usage. By using push_down_predicate when reading from the Data Catalog, the job filters at the source. Other options either do not reduce data read (bookmarks, more DPUs) or are not sufficient alone (Parquet conversion).
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the dataset to Parquet format before processing.
Why it's wrong here
Converting to Parquet improves compression and query performance, but it does not automatically filter partitions. The job would still read all partitions unless a filter is applied. While Parquet is columnar and can reduce I/O, the primary issue is reading unnecessary partitions. The engineer should implement partition pruning, not just change format.
- ✗
Increase the number of DPUs to speed up the job.
Why it's wrong here
Increasing DPUs adds more compute resources but does not reduce the amount of data read. The job would still process all partitions, just faster. This increases cost rather than reducing it. The goal is to process only required partitions, which is achieved by predicate pushdown or partition pruning, not by adding more resources.
- ✗
Enable AWS Glue job bookmarks to track processed partitions.
Why it's wrong here
Job bookmarks help track previously processed data to avoid reprocessing, but they do not filter partitions based on query predicates. They are useful for incremental processing, not for optimizing a job that reads all partitions and filters. The job would still read all partitions on the first run. Bookmarks do not reduce the data scanned for a given run; they only prevent reprocessing of old data.
- ✓
Use predicate pushdown in the Glue job by specifying a push_down_predicate when reading from the AWS Glue Data Catalog.
Why this is correct
Predicate pushdown allows the Glue job to filter partitions at the source based on the push_down_predicate parameter. When reading from the Data Catalog, you can specify a condition like 'year=2023 and month=01' to read only those partitions. This reduces the amount of data read and processed, improving performance and reducing DPU consumption. It is the recommended way to optimize partition pruning in Glue.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.