Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes 500 GB of data. The engineer notices that the job takes several hours and wants to optimize performance. The data is stored in Parquet format and partitioned by date. Which optimization should the engineer implement to improve the job's performance?

⚠ Common exam trap

The trap here is assuming that simply adding more DPUs will solve performance problems, when actually data layout optimizations like predicate pushdown and partition pruning often yield greater benefits for partitioned Parquet data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use predicate pushdown to filter data at the source and partition pruning to read only necessary partitions.

Predicate pushdown and partition pruning are key optimizations for AWS Glue jobs reading partitioned Parquet data. They minimize the amount of data scanned by pushing filters to the source and skipping irrelevant partitions. This reduces I/O and compute time, directly addressing the performance issue without unnecessary cost increases.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the Parquet data to CSV to improve read performance.

    Why it's wrong here

    Parquet is a columnar format optimized for analytics, offering better compression and faster reads than CSV. Converting to CSV would degrade performance and increase storage size. CSV is row-based and does not support predicate pushdown as efficiently. This change would likely slow down the job and increase costs.

  • ✓

    Use predicate pushdown to filter data at the source and partition pruning to read only necessary partitions.

    Why this is correct

    Predicate pushdown and partition pruning allow Glue to read only the relevant partitions and rows from S3, reducing I/O and the amount of data processed. Since the data is partitioned by date, the job can target specific partitions. This significantly speeds up the job and lowers cost. It is a best practice for large datasets in Parquet.

  • ✗

    Increase the number of DPUs for the Glue job to scale horizontally.

    Why it's wrong here

    Increasing DPUs can improve performance by adding more compute resources, but it is not the most effective optimization for a job reading partitioned Parquet data. Without addressing data layout and pushdown, simply adding DPUs may not yield linear speedup and increases cost. The job may still be I/O bound due to inefficient reads.

  • ✗

    Enable job bookmarks to track processed data and avoid reprocessing.

    Why it's wrong here

    Job bookmarks help avoid reprocessing old data, which is useful for incremental loads, but the scenario describes a daily job that processes 500 GB, likely processing new partitions. Bookmarks do not optimize the transformation of that day's data. They reduce redundant work across runs, not the performance of a single run. The job still needs to read and transform all new data.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.