Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue to read from an Amazon S3 bucket that contains data in Apache Parquet format, partitioned by year/month/day. The Glue job needs to read only the data for the last 7 days. The engineer wants to minimize the amount of data scanned and improve job performance. Which approach should be used to filter the partitions efficiently?

⚠ Common exam trap

The trap here is assuming that filtering after reading the data is equivalent to filtering at the source, when in fact pushdown predicates are required to avoid scanning all partitions.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a pushdown predicate in the Glue job's 'create_dynamic_frame.from_catalog' call to filter on partition columns.

Using a pushdown predicate in the 'create_dynamic_frame.from_catalog' call allows AWS Glue to filter partitions at the source, so only the relevant S3 partitions for the last 7 days are read. This minimizes data scanned and improves job performance by avoiding loading unnecessary data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create a separate Glue Data Catalog table for each day's data and read from the last 7 tables.

    Why it's wrong here

    Creating separate tables for each day is not scalable and complicates catalog management. It would require dynamic table creation and additional logic to determine the last 7 days. This approach does not leverage Glue's built-in partition pruning and is not a recommended best practice.

  • ✗

    Read the entire dataset into a DynamicFrame and then apply a filter transformation on the partition columns.

    Why it's wrong here

    Reading the entire dataset first would scan all data in S3, including partitions outside the last 7 days, leading to unnecessary data transfer and processing. Applying a filter afterward does not reduce the data scanned from S3, as the filtering happens after the data is loaded into memory. This approach is inefficient and increases cost and runtime.

  • ✓

    Use a pushdown predicate in the Glue job's 'create_dynamic_frame.from_catalog' call to filter on partition columns.

    Why this is correct

    AWS Glue supports pushdown predicates when reading from the AWS Glue Data Catalog. By specifying a filter expression on partition columns (e.g., 'year >= 2023 AND month = 10 AND day BETWEEN 1 AND 7'), Glue pushes the filter down to the data source, so only the relevant partitions are read from S3. This significantly reduces the amount of data scanned and improves performance.

  • ✗

    Use the 'glueContext.read_from_options' with a 'filter' parameter to specify the partition range.

    Why it's wrong here

    'read_from_options' does not have a 'filter' parameter for partition pruning. While you can specify options like 'paths' to read specific partitions, that requires manually constructing the list of paths, which is not dynamic and error-prone. The correct method is to use a pushdown predicate with the catalog source.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.