DEA-C01 Data Operations and Support Practice Question
A data engineer is running an Amazon EMR cluster with Spark to process log files. The cluster uses instance fleets with m5.xlarge core nodes. The engineer observes that the Spark job is running slower than expected. CloudWatch metrics show that the cluster's CPU utilization is below 20% but memory utilization is near 90%. Which configuration change would most likely improve performance?
⚠ Common exam trap
The trap here is assuming that 'slower than expected' always means insufficient compute, so candidates add nodes or partitions instead of reading the CloudWatch signal that memory — not CPU — is the saturated resource.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use memory-optimized instances (r5.xlarge) for core nodes.
The CloudWatch metrics show CPU below 20% while memory is near 90%, which is the classic signature of a memory-bound Spark workload. Spark executors on m5.xlarge (16 GiB RAM) are spilling to disk or GC-thrashing because the working set exceeds available heap. Switching core nodes to r5.xlarge (memory-optimized, 32 GiB RAM) doubles the memory per node, reducing spills and GC pressure, which directly addresses the bottleneck.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use memory-optimized instances (r5.xlarge) for core nodes.
Why this is correct
Memory-optimised r5.xlarge instances provide a higher memory-to-vCPU ratio than m5.xlarge, directly relieving the near-90% memory utilisation that is throttling Spark executors. Since CPU sits below 20%, the bottleneck is memory capacity, not compute, so swapping core nodes to r5.xlarge lets executors hold larger partitions without spilling to disk.
- ✗
Increase the number of core nodes from 5 to 10.
Why it's wrong here
Adding core nodes scales CPU and disk capacity, but the bottleneck is memory: each executor's heap is exhausted, causing spills and GC pauses. More nodes with the same m5.xlarge profile cannot relieve per-node memory pressure. Horizontal scaling suits CPU-bound or storage-bound jobs, not memory-constrained executors.
- ✗
Increase the number of Spark shuffle partitions.
Why it's wrong here
Shuffle partitions govern task parallelism, not executor memory; with memory near 90% and CPU idle, raising partition count adds overhead without relieving the bottleneck. It tempts because partition tuning is a common Spark performance lever, but it targets skew and parallelism rather than the memory pressure shown here.
- ✗
Decrease the number of core nodes to reduce overhead.
Why it's wrong here
Removing core nodes reduces aggregate memory further, worsening spills and garbage collection on the already-saturated executors. Fewer nodes cannot lower per-executor heap demand. Shrinking a cluster suits over-provisioned, CPU-idle workloads where cost reduction matters, not a memory-bound Spark job.
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.