DEA-C01 Data Operations and Support Practice Question
A data engineer is troubleshooting an AWS Glue ETL job that fails with the error 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large number of small files in Amazon S3. Which action would MOST effectively resolve the issue?
⚠ Common exam trap
DEA-C01 often tests the misconception that scaling up worker type or count solves OutOfMemory errors, when the root cause (many small files) requires the groupFiles optimization instead.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable S3 groupFiles option in the Glue job
The 'java.lang.OutOfMemoryError: Java heap space' in AWS Glue when processing many small files is typically caused by the driver or executor accumulating too many file metadata objects. Enabling the S3 groupFiles option (with groupSize and groupFiles parameters) consolidates small files into larger groups, reducing the number of objects processed and alleviating heap pressure. This directly addresses the root cause of the memory issue.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable S3 groupFiles option in the Glue job
Why this is correct
Grouping coalesces many small S3 objects into larger input partitions, so Glue reads far fewer files and holds less per-file metadata and buffer overhead in the driver and executors. This directly relieves the Java heap exhaustion caused by the large number of small files described in the stem.
- ✗
Change the worker type to G.1X
Why it's wrong here
G.1X workers offer the same 16 GB memory as the default G.1X already in use, so heap pressure from many small files persists; the fix is G.2X or grouping files. G.1X suits memory-light, CPU-bound jobs, not this out-of-memory scenario.
- ✗
Increase the number of workers in the Glue job
Why it's wrong here
Does not reduce per-worker memory pressure.
- ✗
Use a G.2X worker type with more memory
Why it's wrong here
G.2X adds memory per worker, but the heap error stems from too many small files being listed and read concurrently; compaction or grouping into larger objects reduces task overhead. G.2X is correct when per-worker memory is genuinely the constraint, not file-count-driven heap pressure.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.