DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to read from an Amazon S3 bucket that contains data in Apache Parquet format, partitioned by year/month/day. The Glue job needs to read only the data for the last 7 days. The engineer wants to minimize the amount of data scanned and improve job performance. Which approach should be used to filter the partitions efficiently?
⚠ Common exam trap
The trap here is assuming that filtering after reading the data is equivalent to filtering at the source, when in fact pushdown predicates are required to avoid scanning all partitions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a pushdown predicate in the Glue job's 'create_dynamic_frame.from_catalog' call to filter on partition columns.
Using a pushdown predicate in the 'create_dynamic_frame.from_catalog' call allows AWS Glue to filter partitions at the source, so only the relevant S3 partitions for the last 7 days are read. This minimizes data scanned and improves job performance by avoiding loading unnecessary data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a separate Glue Data Catalog table for each day's data and read from the last 7 tables.
Why it's wrong here
Creating separate tables for each day is not scalable and complicates catalog management. It would require dynamic table creation and additional logic to determine the last 7 days. This approach does not leverage Glue's built-in partition pruning and is not a recommended best practice.
- ✗
Read the entire dataset into a DynamicFrame and then apply a filter transformation on the partition columns.
Why it's wrong here
Reading the entire dataset first would scan all data in S3, including partitions outside the last 7 days, leading to unnecessary data transfer and processing. Applying a filter afterward does not reduce the data scanned from S3, as the filtering happens after the data is loaded into memory. This approach is inefficient and increases cost and runtime.
- ✓
Use a pushdown predicate in the Glue job's 'create_dynamic_frame.from_catalog' call to filter on partition columns.
Why this is correct
AWS Glue supports pushdown predicates when reading from the AWS Glue Data Catalog. By specifying a filter expression on partition columns (e.g., 'year >= 2023 AND month = 10 AND day BETWEEN 1 AND 7'), Glue pushes the filter down to the data source, so only the relevant partitions are read from S3. This significantly reduces the amount of data scanned and improves performance.
- ✗
Use the 'glueContext.read_from_options' with a 'filter' parameter to specify the partition range.
Why it's wrong here
'read_from_options' does not have a 'filter' parameter for partition pruning. While you can specify options like 'paths' to read specific partitions, that requires manually constructing the list of paths, which is not dynamic and error-prone. The correct method is to use a pushdown predicate with the catalog source.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.