Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is building an AWS Glue job that reads from a large Parquet dataset in Amazon S3 partitioned by year/month/day and writes aggregated results to Amazon Redshift. The job currently reads all partitions and takes several hours. The engineer wants the job to process only partitions from the last seven days and reduce runtime. Which change should the engineer make?

⚠ Common exam trap

The trap here is reaching for more DPUs or bookmarks, when the actual bottleneck is scanning partitions that fall outside the desired date range.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Pass a pushdown predicate in the from_catalog call that filters on the year, month, and day partition columns for the last seven days.

Partition pruning is the key optimization for partitioned datasets on S3. Supplying a pushdown predicate that references the year, month, and day partition columns lets the Glue reader skip non-matching partitions entirely, so only seven days of data are listed and read. More workers, format conversion, and bookmarks do not restrict the read window and therefore do not solve the runtime problem.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Pass a pushdown predicate in the from_catalog call that filters on the year, month, and day partition columns for the last seven days.

    Why this is correct

    A pushdown predicate on partition columns is evaluated by the Glue Data Catalog and S3 reader before data is loaded, so only the matching partitions are listed and read. Restricting to the last seven days dramatically reduces the number of files scanned, cutting runtime and cost while satisfying the requirement.

  • ✗

    Convert the source data to ORC format to reduce the bytes read per partition.

    Why it's wrong here

    ORC is columnar and can reduce I/O versus row formats, but the dataset is already Parquet, which is also columnar. Converting formats adds a rewrite step and does not restrict the read to the last seven days, so the job would still scan the entire historical dataset and remain slow.

  • ✗

    Increase the number of DPUs allocated to the Glue job and enable auto scaling.

    Why it's wrong here

    Adding DPUs and enabling auto scaling increases parallel compute, which can shorten runtime for CPU-bound work, but the job still reads every partition in the dataset. The dominant cost here is scanning unnecessary data, so more workers do not address the root cause and would raise cost without meeting the seven-day requirement.

  • ✗

    Enable job bookmarks so the job processes only new files since the last successful run.

    Why it's wrong here

    Bookmarks track which files have been processed across runs, which is useful for incremental ingestion, but they do not filter by partition date. If the job has never run before, bookmarks provide no reduction; and if historical files remain unprocessed, they would still be read, so the seven-day window would not be honored.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.