DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is building an AWS Glue job that reads from a large Parquet dataset in Amazon S3 partitioned by year/month/day and writes aggregated results to Amazon Redshift. The job currently reads all partitions and takes several hours. The engineer wants the job to process only partitions from the last seven days and reduce runtime. Which change should the engineer make?
⚠ Common exam trap
The trap here is reaching for more DPUs or bookmarks, when the actual bottleneck is scanning partitions that fall outside the desired date range.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pass a pushdown predicate in the from_catalog call that filters on the year, month, and day partition columns for the last seven days.
Partition pruning is the key optimization for partitioned datasets on S3. Supplying a pushdown predicate that references the year, month, and day partition columns lets the Glue reader skip non-matching partitions entirely, so only seven days of data are listed and read. More workers, format conversion, and bookmarks do not restrict the read window and therefore do not solve the runtime problem.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Pass a pushdown predicate in the from_catalog call that filters on the year, month, and day partition columns for the last seven days.
Why this is correct
A pushdown predicate on partition columns is evaluated by the Glue Data Catalog and S3 reader before data is loaded, so only the matching partitions are listed and read. Restricting to the last seven days dramatically reduces the number of files scanned, cutting runtime and cost while satisfying the requirement.
- ✗
Convert the source data to ORC format to reduce the bytes read per partition.
Why it's wrong here
ORC is columnar and can reduce I/O versus row formats, but the dataset is already Parquet, which is also columnar. Converting formats adds a rewrite step and does not restrict the read to the last seven days, so the job would still scan the entire historical dataset and remain slow.
- ✗
Increase the number of DPUs allocated to the Glue job and enable auto scaling.
Why it's wrong here
Adding DPUs and enabling auto scaling increases parallel compute, which can shorten runtime for CPU-bound work, but the job still reads every partition in the dataset. The dominant cost here is scanning unnecessary data, so more workers do not address the root cause and would raise cost without meeting the seven-day requirement.
- ✗
Enable job bookmarks so the job processes only new files since the last successful run.
Why it's wrong here
Bookmarks track which files have been processed across runs, which is useful for incremental ingestion, but they do not filter by partition date. If the job has never run before, bookmarks provide no reduction; and if historical files remain unprocessed, they would still be read, so the seven-day window would not be honored.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.