MLS-C01 Data Engineering Practice Question
A company uses Amazon EMR to run Spark jobs on a cluster with 10 core nodes of type r5.xlarge. The jobs are I/O intensive and read large amounts of data from S3. The team notices high network throughput but low CPU utilization. Which configuration change would improve job performance at the same cost?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Replace the core nodes with r5d.xlarge instances that have local SSDs.
R5d instances include local NVMe SSDs. These SSDs can be used for caching intermediate data during Spark jobs, reducing the need to read from and write to S3 over the network. This directly addresses the I/O bottleneck and high network throughput observed, improving job performance. Option A is incorrect because moving to general-purpose m5 instances does not provide local SSDs and may not improve I/O. Option B is incorrect because doubling the number of core nodes would increase cost significantly without necessarily solving the I/O issue. Option D is incorrect because spot instances reduce cost but do not inherently improve I/O performance; they may even add instability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Change the instance type to m5.xlarge (general purpose) to balance resources.
Why it's wrong here
m5 instances have lower network and EBS performance than r5.
- ✗
Increase the number of core nodes to 20.
Why it's wrong here
This doubles the cost.
- ✓
Replace the core nodes with r5d.xlarge instances that have local SSDs.
Why this is correct
Local SSDs provide high I/O for caching, reducing network traffic.
- ✗
Use spot instances for the core nodes to save cost and reinvest in more nodes.
Why it's wrong here
Spot instances reduce cost but do not improve performance directly.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLS-C01 question from scratch — 1,672 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.