Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue Studio to create a visual ETL job that reads from an Amazon S3 bucket containing JSON files, applies a filter transformation, and writes the output to Amazon Redshift. The job must run daily. The engineer notices that the job is taking a long time to complete and wants to improve performance. Which action should the engineer take to optimize the job?

⚠ Common exam trap

The trap here is assuming that adding more DPUs is always the first step to improve AWS Glue job performance, when data format optimization often yields greater benefits at lower cost.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the source JSON files to Parquet format before running the ETL job.

Converting JSON to Parquet improves ETL performance because Parquet is columnar, compressed, and requires less I/O and CPU to parse. AWS Glue jobs benefit significantly from columnar formats when reading from S3. While increasing DPUs or enabling bookmarks can help in some cases, the most direct and effective optimization for slow JSON processing is to use a more efficient file format.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Convert the source JSON files to Parquet format before running the ETL job.

    Why this is correct

    Parquet is a columnar format that offers better compression and faster read performance compared to JSON. Converting the source data to Parquet reduces I/O and CPU overhead during the ETL job, significantly improving performance. This is a common best practice for AWS Glue jobs reading from S3, especially when the data is large. It directly addresses the slow read and transform steps.

  • ✗

    Use a larger number of smaller files instead of fewer large files.

    Why it's wrong here

    Having many small files can actually degrade performance because it increases the number of read operations and metadata overhead. AWS Glue performs better with fewer, larger files. Therefore, this action would likely worsen performance rather than improve it. The recommendation is to compact small files, not create more of them.

  • ✗

    Increase the number of AWS Glue DPUs allocated to the job.

    Why it's wrong here

    Increasing DPUs can improve performance for CPU-intensive or memory-intensive jobs, but the primary bottleneck here is likely data format and partitioning. JSON is not columnar and parsing it is slower than Parquet. Simply adding more DPUs may not address the inefficiency of reading and transforming JSON, and it increases cost. It is not the most effective optimization for this scenario.

  • ✗

    Enable job bookmarks to track processed files and avoid reprocessing.

    Why it's wrong here

    Job bookmarks help avoid reprocessing already processed data, which can reduce runtime if the job is reprocessing old files. However, in a daily job that reads new data each day, bookmarks may not provide significant performance gains unless the job is mistakenly reprocessing all historical data. The main performance issue is likely due to the inefficient JSON format, not reprocessing.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.