mediumMultiple Choice
MLA-C01 Practice Question: Refer to the exhibit
Exhibit
{
"PolicyARN": "arn:aws:autoscaling:us-east-1:123456789012:scalingPolicy:policy-1",
"PolicyName": "SageMakerEndpointScalingPolicy",
"PolicyType": "TargetTrackingScaling",
"TargetTrackingScalingPolicyConfiguration": {
"TargetValue": 70.0,
"PredefinedMetricSpecification": {
"PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"
},
"ScaleInCooldown": 600,
"ScaleOutCooldown": 200
}
}Refer to the exhibit. A team observes that their SageMaker endpoint scales out quickly when load increases, but scales in very slowly when load decreases, causing over-provisioning. What is the most likely cause?
⚠ Common exam trap
A common mix-up: candidates confuse cooldown periods with scaling thresholds, assuming that slow scale-in is caused by a high TargetValue or wrong metric, rather than recognizing that cooldown timers directly control the delay between scaling actions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ScaleInCooldown is too high
A high ScaleInCooldown value causes the SageMaker endpoint to wait too long before initiating a scale-in event after load decreases. This delay prevents the endpoint from releasing resources promptly, leading to over-provisioning. In contrast, the scaling out behavior is unaffected by this cooldown, which explains why the endpoint scales out quickly but scales in slowly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
TargetValue is too high
Why it's wrong here
TargetValue sets the metric threshold that triggers scaling; raising it makes the endpoint tolerate higher load before adding instances, affecting scale-out rather than the slow scale-in. It tempts because a high threshold suggests over-provisioning, yet the delayed decrease is caused by ScaleInCooldown, not TargetValue.
- ✗
ScaleOutCooldown is too low
Why it's wrong here
ScaleOutCooldown governs how long after a scale-out event before another scaling action; a low value speeds further scale-out, not scale-in. It tempts because cooldown parameters sound responsible for sluggish scaling, but the slow decrease stems from ScaleInCooldown being too high, delaying removal of idle instances.
- ✓
ScaleInCooldown is too high
Why this is correct
ScaleInCooldown specifies the seconds before another scale-in action after one completes. An excessively high value delays each scale-in step, so the endpoint removes instances far too slowly when load drops, producing sustained over-provisioning despite rapid scale-out.
- ✗
Wrong predefined metric selected
Why it's wrong here
Wrong predefined metric cannot explain asymmetric scaling: scale-out and scale-in are governed by separate CloudWatch alarms, so slow scale-in points to a long scale-in cooldown or missing scale-in policy. A predefined metric is chosen when built-in metrics such as CPUUtilization or InvocationsPerInstance already match the workload.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.