MLS-C01 Data Engineering Practice Question
A company uses AWS Glue ETL jobs to transform data from Amazon RDS for MySQL to Amazon S3. The transformation includes aggregations and joins. The job runs daily and processes approximately 100 GB of data. Recently, the job started failing with memory errors on the worker nodes. Which approach would MOST effectively resolve the issue without changing the logic?
⚠ Common exam trap
The trap here is that candidates might confuse worker type (memory per worker) with number of workers (total cluster memory), incorrectly assuming a larger worker type always helps, when in fact increasing the number of workers is the direct fix for memory errors in distributed Spark jobs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of workers in the job configuration
Increasing the number of workers distributes the memory load across more nodes, which directly addresses memory errors in a Spark ETL job without altering the transformation logic. AWS Glue Spark jobs process data in memory across workers, and insufficient total memory causes out-of-memory errors when handling 100 GB of data with aggregations and joins.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch from a Spark ETL job to a Python shell job
Why it's wrong here
Python shell jobs are not designed for large-scale data processing.
- ✗
Decrease the number of workers to reduce overhead
Why it's wrong here
Fewer workers increase memory pressure per worker.
- ✗
Change the worker type from G.2X to G.1X to increase memory per worker
Why it's wrong here
G.1X has less memory than G.2X, not more.
- ✓
Increase the number of workers in the job configuration
Why this is correct
More workers distribute the data processing, reducing memory per node.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This MLS-C01 question is part of Courseiva's 1,672-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.