Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue to transform data from an Amazon S3 bucket. The Glue job reads JSON files, applies complex transformations, and writes the output to another S3 bucket in Parquet format. The job runs daily and must complete within a 2-hour window. The engineer notices that the job is taking longer than expected and wants to optimize performance. The source data is partitioned by date, and the job uses a dynamic frame. Which optimization should the engineer implement to improve performance?

⚠ Common exam trap

The trap here is assuming that adding more DPUs is always the best way to speed up a Glue job, when in fact reducing the amount of data read via pushdown predicates is often more effective and cost-efficient.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use pushdown predicates to filter data at the source based on partition columns.

Pushdown predicates enable AWS Glue to filter data at the source based on partition columns, significantly reducing the amount of data read from S3. Since the data is partitioned by date, specifying a predicate for the desired date range ensures only relevant partitions are processed, leading to faster job completion and lower cost. This is a targeted optimization that addresses the performance bottleneck without unnecessary resource scaling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of DPUs allocated to the Glue job.

    Why it's wrong here

    Increasing DPUs can improve performance by adding more compute resources, but it may not address the root cause if the job is inefficient due to data skew or excessive small files. Simply adding DPUs increases cost and might not yield proportional speedup. The engineer should first optimize the data layout and transformation logic before scaling horizontally.

  • ✗

    Enable job bookmarks to track processed data and avoid reprocessing.

    Why it's wrong here

    Job bookmarks help prevent reprocessing of already processed data, which is useful for incremental loads, but the scenario states the job runs daily on partitioned data. If the job is already processing only new partitions, bookmarks may not significantly improve performance. The issue is likely within the job's execution, not data reprocessing.

  • ✗

    Convert the dynamic frame to a Spark DataFrame and use Spark SQL for transformations.

    Why it's wrong here

    Converting to a DataFrame can provide more optimization opportunities through Spark's Catalyst optimizer, but it requires rewriting transformations and may not yield immediate gains. It is not a guaranteed performance fix and adds complexity. The primary issue is likely reading unnecessary data, which pushdown predicates solve directly without code changes.

  • ✓

    Use pushdown predicates to filter data at the source based on partition columns.

    Why this is correct

    Pushdown predicates allow Glue to filter data at the source, reducing the amount of data read and processed. Since the data is partitioned by date, the engineer can specify a predicate to read only the relevant partitions. This minimizes I/O and speeds up the job, often more effectively than adding resources. It directly addresses the performance bottleneck by reducing data volume.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.