Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A data engineer is configuring an AWS Glue ETL job to read from an Amazon S3 bucket that contains Apache Parquet files partitioned by year, month, and day. The engineer wants the job to only process data for the year 2023 and month 10, and to minimize the amount of data scanned. The Glue job uses the Glue Data Catalog table `sales_data` with the correct partition structure. What is the MOST efficient way to configure the job to read only the required partitions?

⚠ Common exam trap

The trap here is assuming that filtering after reading the data is equivalent to partition pruning, but it still scans all data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Glue `create_dynamic_frame.from_catalog` with `push_down_predicate` set to `year=='2023' and month=='10'`.

The correct approach is to use `push_down_predicate` with `create_dynamic_frame.from_catalog`. This pushes the filter down to the Glue Data Catalog, so only the relevant partitions are read from Amazon S3. It minimizes data scanned, reduces cost, and improves job performance. Other methods either read all data first or bypass the catalog, which is less efficient and harder to maintain.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Read the entire `sales_data` table into a DynamicFrame and then apply a `Filter` transform with the condition `year=='2023' and month=='10'`.

    Why it's wrong here

    Reading the entire table first loads all partitions into memory, which scans all data in S3, increasing cost and runtime. The Filter transform is applied after the data is read, so it does not reduce the amount of data scanned. For large datasets, this approach is inefficient and defeats the purpose of partitioning. It also may cause out-of-memory errors if the dataset is large.

  • ✗

    Use the Glue `create_dynamic_frame.from_options` with the S3 path `s3://bucket/sales_data/year=2023/month=10/` and format `parquet`.

    Why it's wrong here

    While this reads only the specified prefix, it bypasses the Glue Data Catalog and does not automatically handle partition columns or schema. It also requires hardcoding the path, which is less flexible. Additionally, if the table has other partition columns like day, this method would read all days under the month, but it might not correctly infer partition values. Using the catalog with pushdown is more maintainable and efficient.

  • ✓

    Use the Glue `create_dynamic_frame.from_catalog` with `push_down_predicate` set to `year=='2023' and month=='10'`.

    Why this is correct

    Using `push_down_predicate` with `create_dynamic_frame.from_catalog` allows Glue to filter partitions at the catalog level, so only the specified partitions are read from S3. This reduces data scanned and improves performance. The predicate syntax uses SQL-like expressions on partition columns, and Glue leverages the partition metadata to avoid listing and reading unnecessary partitions. This is the recommended approach for partitioned data in Glue ETL jobs.

  • ✗

    Create a new Glue Data Catalog table that points only to the `year=2023/month=10` prefix, and then read from that table.

    Why it's wrong here

    Creating a new table for a specific partition is not scalable and duplicates metadata. It also requires manual updates when new partitions are added. This approach does not leverage the existing partition structure and can lead to inconsistencies. The original table already contains the partition metadata, so using pushdown predicates is simpler and more efficient. This method is not recommended for ad-hoc or recurring ETL jobs.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.