Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A company uses AWS Glue ETL jobs to transform data in S3. The job runs successfully but takes longer than expected. The data is in Parquet format and partitioned by date. Which change would most improve performance without increasing cost?

⚠ Common exam trap

A common mix-up: candidates assume performance issues are solved by adding more resources (DPUs) or changing file formats, when the real bottleneck is reading unnecessary data due to lack of partition pruning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable pushdown predicates to filter partitions early.

Pushdown predicates allow AWS Glue to filter data at the storage layer (e.g., S3 partition pruning) before reading it into memory. Since the data is partitioned by date, enabling pushdown predicates reduces the amount of data scanned, which directly decreases job runtime without requiring additional DPUs or changing the data format.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Repartition the data by a different column.

    Why it's wrong here

    Repartitioning by another column removes the date-based partition pruning that limits which S3 objects Glue reads, forcing broader scans. It is tempting because repartitioning can fix skewed partitions, and it would be correct if a single date partition held disproportionate data causing straggler tasks.

  • ✗

    Convert Parquet to CSV for faster serialization.

    Why it's wrong here

    CSV is row-oriented and uncompressed text, so Glue reads and parses far more bytes than columnar Parquet, worsening runtime. It is tempting because CSV is human-readable and simple to generate, and it would be correct only when downstream tools cannot read Parquet.

  • ✗

    Increase the number of DPUs for the job.

    Why it's wrong here

    Adding DPUs raises the job's compute cost, violating the no-cost-increase requirement, and cannot fix inefficient partition pruning. It is tempting because more workers speed up embarrassingly parallel stages, and it would be correct if the bottleneck were genuinely CPU-bound with no cheaper tuning available.

  • ✓

    Enable pushdown predicates to filter partitions early.

    Why this is correct

    Pushdown predicates let AWS Glue push partition and column filters to the S3 data source, so only matching partitions and row groups are read. With date partitioning, this prunes irrelevant partitions early, cutting I/O and runtime without adding compute capacity or cost.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.