DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue Studio to create a visual ETL job that reads from an Amazon S3 bucket containing JSON files, applies a filter transformation, and writes the output to Amazon Redshift. The job must run daily. The engineer notices that the job is taking a long time to complete and wants to improve performance. Which action should the engineer take to optimize the job?
⚠ Common exam trap
The trap here is assuming that adding more DPUs is always the first step to improve AWS Glue job performance, when data format optimization often yields greater benefits at lower cost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the source JSON files to Parquet format before running the ETL job.
Converting JSON to Parquet improves ETL performance because Parquet is columnar, compressed, and requires less I/O and CPU to parse. AWS Glue jobs benefit significantly from columnar formats when reading from S3. While increasing DPUs or enabling bookmarks can help in some cases, the most direct and effective optimization for slow JSON processing is to use a more efficient file format.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Convert the source JSON files to Parquet format before running the ETL job.
Why this is correct
Parquet is a columnar format that offers better compression and faster read performance compared to JSON. Converting the source data to Parquet reduces I/O and CPU overhead during the ETL job, significantly improving performance. This is a common best practice for AWS Glue jobs reading from S3, especially when the data is large. It directly addresses the slow read and transform steps.
- ✗
Use a larger number of smaller files instead of fewer large files.
Why it's wrong here
Having many small files can actually degrade performance because it increases the number of read operations and metadata overhead. AWS Glue performs better with fewer, larger files. Therefore, this action would likely worsen performance rather than improve it. The recommendation is to compact small files, not create more of them.
- ✗
Increase the number of AWS Glue DPUs allocated to the job.
Why it's wrong here
Increasing DPUs can improve performance for CPU-intensive or memory-intensive jobs, but the primary bottleneck here is likely data format and partitioning. JSON is not columnar and parsing it is slower than Parquet. Simply adding more DPUs may not address the inefficiency of reading and transforming JSON, and it increases cost. It is not the most effective optimization for this scenario.
- ✗
Enable job bookmarks to track processed files and avoid reprocessing.
Why it's wrong here
Job bookmarks help avoid reprocessing already processed data, which can reduce runtime if the job is reprocessing old files. However, in a daily job that reads new data each day, bookmarks may not provide significant performance gains unless the job is mistakenly reprocessing all historical data. The main performance issue is likely due to the inefficient JSON format, not reprocessing.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.