Courseiva
Data EngineeringmediumMultiple ChoiceObjective-mapped

MLS-C01 Data Engineering Practice Question

A company uses Amazon EMR to run Spark jobs on a cluster with 10 core nodes of type r5.xlarge. The jobs are I/O intensive and read large amounts of data from S3. The team notices high network throughput but low CPU utilization. Which configuration change would improve job performance at the same cost?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Replace the core nodes with r5d.xlarge instances that have local SSDs.

R5d instances include local NVMe SSDs. These SSDs can be used for caching intermediate data during Spark jobs, reducing the need to read from and write to S3 over the network. This directly addresses the I/O bottleneck and high network throughput observed, improving job performance. Option A is incorrect because moving to general-purpose m5 instances does not provide local SSDs and may not improve I/O. Option B is incorrect because doubling the number of core nodes would increase cost significantly without necessarily solving the I/O issue. Option D is incorrect because spot instances reduce cost but do not inherently improve I/O performance; they may even add instability.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Change the instance type to m5.xlarge (general purpose) to balance resources.

    Why it's wrong here

    m5 instances have lower network and EBS performance than r5.

  • Increase the number of core nodes to 20.

    Why it's wrong here

    This doubles the cost.

  • Replace the core nodes with r5d.xlarge instances that have local SSDs.

    Why this is correct

    Local SSDs provide high I/O for caching, reducing network traffic.

  • Use spot instances for the core nodes to save cost and reinvest in more nodes.

    Why it's wrong here

    Spot instances reduce cost but do not improve performance directly.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every MLS-C01 question from scratch — 1,672 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.