hardMultiple Choice
MLA-C01 Practice Question: A company deploys a model using SageMaker…
A company deploys a model using SageMaker real-time endpoint with auto scaling. They observe that during a traffic spike, the endpoint quickly scales up to 10 instances, but after the spike, it takes a long time to scale down, leading to high costs. The scaling policy is based on a simple average CPU utilization threshold. Which adjustment would optimize the scaling down behavior?
⚠ Common exam trap
Candidates often confuse cooldown periods with step adjustments, thinking that larger scale-in steps will speed up the process, when in fact the cooldown period controls the timing of when scaling actions can occur.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decrease the scale-in cooldown period to allow the endpoint to scale down faster when utilization drops.
Decreasing the scale-in cooldown period allows the endpoint to respond more quickly to sustained drops in CPU utilization. By default, SageMaker auto scaling uses cooldown periods to prevent rapid fluctuations; a long scale-in cooldown delays the termination of instances after utilization falls, keeping costs high. Reducing this cooldown lets the endpoint scale down faster when the spike subsides, directly addressing the problem.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the scale-in cooldown period to prevent premature scale-down.
Why it's wrong here
Lengthening the scale-in cooldown deliberately delays instance removal, worsening the cost problem rather than fixing it. Cooldowns exist to stop flapping when metrics oscillate around a threshold, so extending one would be right only if premature scale-down were causing instability, which the scenario does not describe.
- ✓
Decrease the scale-in cooldown period to allow the endpoint to scale down faster when utilization drops.
Why this is correct
Scale-in cooldown governs how long the endpoint waits after a scale-in before removing more instances. Shortening it lets the endpoint shed surplus instances sooner once CPU utilisation drops, cutting idle capacity costs after the traffic spike.
- ✗
Use a step scaling policy with a larger step adjustment for scale-in.
Why it's wrong here
Step scaling with larger scale-in adjustments removes instances faster once the threshold is breached, but the delay stems from the policy's default cooldown and evaluation periods, which this does not shorten. Step scaling suits workloads with sharply tiered, predictable demand rather than smoothing a gradual post-spike decline.
- ✗
Change the scaling policy to use memory utilization instead of CPU.
Why it's wrong here
Memory utilisation does not govern when instances are removed; the policy still waits for its cooldown and evaluation windows before scaling in, so the delay persists. Memory-based policies suit memory-bound inference workloads where CPU stays low, not this CPU-driven endpoint.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.