DEA-C01 Data Operations and Support Practice Question
A company is running an Amazon EMR cluster with Spark for data processing. The data engineer wants to automatically scale the core and task nodes based on the YARN memory and CPU utilization. Which scaling metric should the engineer use for the EMR managed scaling policy?
⚠ Common exam trap
The trap is picking a familiar CloudWatch metric like CPUUtilization or DiskIOPS instead of the YARN-centric metric that EMR managed scaling actually uses; candidates who have not read the EMR managed scaling documentation often fall for this.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
YARNMemoryAvailablePercentage
EMR managed scaling policies are designed around YARN metrics, and YARNMemoryAvailablePercentage is the primary metric used to determine when to add or remove core and task nodes. When available YARN memory drops below a threshold, EMR scales out; when it rises above a threshold, EMR scales in. This directly reflects the resource pressure that Spark executors place on the cluster.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
YARNMemoryAvailablePercentage
Why this is correct
YARNMemoryAvailablePercentage reflects the memory available to YARN containers, which governs whether Spark executors can be scheduled on core and task nodes. Scaling on this metric adds or removes nodes in step with actual workload demand, matching the stem's YARN memory and CPU utilisation requirement.
- ✗
CPUUtilization
Why it's wrong here
CPUUtilization reflects processor load only; Spark executors are typically constrained by YARN container memory, so CPU may stay low while memory pressure demands more nodes. It tempts because CPU is the default metric for generic EC2 Auto Scaling, but EMR managed scaling expects YARNMemoryAvailablePercentage for Spark workloads.
- ✗
DiskIOPS
Why it's wrong here
DiskIOPS measures EBS volume throughput, which EMR managed scaling does not evaluate when sizing core and task nodes; Spark memory-bound workloads can idle on I/O while YARN queues stall. It tempts as a storage-bottleneck signal, but that metric suits provisioned-IOPS tuning, not cluster node scaling.
- ✗
HDFSUtilization
Why it's wrong here
HDFSUtilization tracks HDFS storage consumption, which is irrelevant to Spark's in-memory shuffle and executor demand; a cluster can exhaust YARN memory while HDFS sits half empty. It tempts because HDFS capacity matters for Hadoop-centric clusters, but EMR managed scaling uses YARNMemoryAvailablePercentage, not storage.
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.