MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security
A team deploys a SageMaker real-time endpoint and configures it with an auto scaling policy targeting a variant. During a flash sale, traffic spikes and the team notices that the number of instances increases, but the average model latency still climbs above the target. The team wants the scaling behavior to react faster to sudden bursts without over-provisioning during steady periods. Which change should they make to the scaling policy?
⚠ Common exam trap
The trap here is assuming a longer cooldown or a higher metric target improves burst handling, when both actually delay scale-out and let latency rise during sudden spikes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Lower the target value of the predefined InvocationsPerInstance metric and shorten the scale-out cooldown so capacity is added sooner.
Target tracking keeps a chosen metric near its target. Lowering the InvocationsPerInstance target causes scale-out to trigger at a lower request rate per instance, so additional capacity arrives earlier during a burst. Shortening the scale-out cooldown lets successive scale-out actions fire sooner. Scaling in still occurs during quiet periods, so steady-state cost is not inflated.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch the policy from target tracking to a step scaling policy with a larger scale-out cooldown and a smaller scale-in cooldown.
Why it's wrong here
Step scaling can add capacity in tiers, but increasing the scale-out cooldown delays subsequent scale-out actions, which slows the response to a sustained burst rather than speeding it up. Cooldowns exist to prevent thrashing; lengthening the scale-out cooldown works against the goal of reacting faster to sudden traffic spikes.
- ✓
Lower the target value of the predefined InvocationsPerInstance metric and shorten the scale-out cooldown so capacity is added sooner.
Why this is correct
Target tracking adjusts capacity to keep the metric near the target. Lowering the InvocationsPerInstance target means the policy scales out at a lower request rate per instance, adding capacity earlier, and shortening the scale-out cooldown allows consecutive scale-out actions to occur more quickly. Together they make the endpoint respond faster to bursts while still scaling in during steady, low-traffic periods.
- ✗
Replace target tracking with a scheduled scaling policy that adds instances at fixed times each day.
Why it's wrong here
Scheduled scaling adds capacity at predetermined times, which suits predictable daily peaks but cannot react to an unplanned flash sale. A sudden burst outside the schedule would still overwhelm the fleet, so latency would climb. Scheduled scaling is best combined with dynamic policies, not used as a replacement when unpredictable spikes are the concern.
- ✗
Increase the target value of the predefined InvocationsPerInstance metric and enable a longer scale-in cooldown to stabilize the fleet.
Why it's wrong here
Raising the InvocationsPerInstance target tells the policy to tolerate more requests per instance before scaling out, which delays capacity additions and worsens latency during a burst. A longer scale-in cooldown only affects how quickly capacity is removed, not how fast it is added, so this change moves away from the desired responsiveness.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.