DEA-C01 Data Operations and Support Practice Question
A data engineer is using AWS Glue to run a PySpark ETL job that processes millions of small JSON files stored in an Amazon S3 bucket. The job is experiencing high memory usage and failing with 'Container killed by YARN for exceeding memory limits'. The engineer wants to optimize the job to handle the data more efficiently without changing the source data format. Which solution will MOST effectively reduce memory usage and improve performance?
⚠ Common exam trap
The trap here is assuming that increasing DPUs or enabling bookmarks will solve the memory issue, when the real problem is the inefficiency of processing many small files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS Glue's groupFiles and groupSize options to combine small files into larger chunks during the read operation.
The groupFiles and groupSize options in AWS Glue are specifically designed to handle large numbers of small files by grouping them into larger chunks during read, which reduces the number of Spark partitions and memory overhead. This optimization directly targets the root cause of the memory failures and improves job efficiency without altering the source data format.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the JSON files to Parquet format using an AWS Glue crawler before running the ETL job.
Why it's wrong here
Converting to Parquet can improve query performance and reduce storage, but it requires an additional transformation step and does not directly solve the memory issue during the ETL job. The crawler only catalogs data; it does not convert formats. Moreover, the conversion itself would need to process the small files, potentially facing the same memory problem.
- ✓
Use AWS Glue's groupFiles and groupSize options to combine small files into larger chunks during the read operation.
Why this is correct
The groupFiles and groupSize options in AWS Glue allow the job to coalesce multiple small files into larger groups, reducing the number of tasks and memory overhead per file. This directly addresses the inefficiency of processing many small files by minimizing the number of Spark partitions and the associated memory pressure, leading to better performance and lower resource consumption.
- ✗
Increase the number of DPUs allocated to the AWS Glue job to provide more memory per executor.
Why it's wrong here
Adding more DPUs increases the total cluster resources and may temporarily alleviate memory pressure, but it does not address the fundamental inefficiency of processing numerous small files. Each file still incurs overhead, and simply scaling up can be cost-prohibitive without resolving the root cause. This is a temporary workaround, not an effective optimization.
- ✗
Enable AWS Glue job bookmarks to track processed files and avoid reprocessing.
Why it's wrong here
Job bookmarks track previously processed data to prevent reprocessing, but they do not reduce the memory footprint of processing millions of small files in a single job run. The high memory usage stems from the overhead of reading many small files, not from reprocessing. Bookmarks are useful for incremental loads but do not address the core performance issue here.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.