DEA-C01 Data Operations and Support Practice Question
A data engineer is optimizing an AWS Glue ETL job that reads from Amazon S3 and writes to Amazon Redshift. The job currently uses a single large file and takes hours to complete. The engineer wants to improve performance by using partitioning and parallelism. Which TWO actions should the engineer take? (Choose two.)
⚠ Common exam trap
The trap here is focusing on incremental processing features like job bookmarks or schema handling with DynamicFrames, which do not directly speed up a single large dataset.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the input data to a partitioned format, such as Parquet, and store it in S3 with a partition key.
To improve performance, the engineer should partition the input data and use a columnar format like Parquet, which enables Glue to read only necessary data and process partitions in parallel. Additionally, increasing the number of DPUs provides more compute resources to handle the parallel tasks. Together, these actions address both data layout and compute capacity, leading to faster job execution.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure the job to write output to a single file to reduce the number of S3 PUT requests.
Why it's wrong here
Writing to a single file reduces the number of output files but can create a bottleneck and limit parallelism during the write phase. It may also cause memory issues if the data is large. This approach contradicts the goal of improving performance through parallelism. It is better to write multiple partitioned files to leverage concurrent writes.
- ✓
Convert the input data to a partitioned format, such as Parquet, and store it in S3 with a partition key.
Why this is correct
Partitioning the data and using a columnar format like Parquet allows Glue to read only relevant partitions and leverage predicate pushdown. This reduces I/O and enables parallel processing across partitions. It directly addresses the slow performance caused by a single large file by splitting data into manageable chunks that can be processed concurrently.
- ✗
Enable job bookmarks to track processed data and avoid reprocessing.
Why it's wrong here
Job bookmarks help with incremental processing by tracking previously processed data, but they do not directly improve performance for a single large file. They reduce redundant work in subsequent runs but do not address parallelism or partitioning for the current dataset. The scenario focuses on speeding up the current job, not incremental processing.
- ✓
Increase the number of DPUs allocated to the Glue job to allow more parallel tasks.
Why this is correct
Increasing DPUs provides more compute resources, which allows Glue to run more concurrent tasks and process data in parallel. This is effective when the job is resource-bound. However, it should be combined with data partitioning to fully utilize the additional capacity. It directly improves performance by scaling the job's execution environment.
- ✗
Use the Glue DynamicFrame instead of a DataFrame to enable automatic schema inference.
Why it's wrong here
DynamicFrame provides schema flexibility and handles inconsistencies, but it does not inherently improve performance for large datasets. In fact, DynamicFrames can have overhead compared to DataFrames. The scenario requires performance optimization through partitioning and parallelism, not schema handling. This choice does not address the root cause of slow processing.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.