MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security
A company wants to reduce costs for a SageMaker real-time endpoint that has variable traffic. Which feature allows the endpoint to automatically adjust instance count based on demand?
⚠ Common exam trap
MLA-C01 often tests the confusion between cost-optimization features — candidates pick Savings Plans (a billing discount) when the question is actually about dynamic capacity adjustment via autoscaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Auto Scaling for SageMaker endpoints
Auto Scaling for SageMaker endpoints is the native capability that dynamically adjusts the number of instances behind a real-time endpoint based on CloudWatch metrics such as InvocationsPerInstance or ModelLatency. It uses Application Auto Scaling policies (target tracking or step scaling) to add instances during traffic spikes and remove them during lulls, directly reducing cost for variable workloads. Savings Plans and Inference Recommender do not perform runtime scaling.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
SageMaker Savings Plans
Why it's wrong here
Savings Plans are a billing commitment giving discounted rates on compute usage; they do not change instance count in response to traffic. It is tempting because they genuinely reduce cost for predictable, steady workloads, but the scenario requires automatic scaling, which Savings Plans cannot perform.
- ✗
SageMaker Inference Recommender
Why it's wrong here
Inference Recommender analyses workloads and suggests instance types and configurations; it does not scale a live endpoint's instance count. It is tempting because it optimises instance selection for cost and performance, which is correct during initial deployment planning, not for automatic runtime scaling under variable demand.
- ✗
SageMaker Model Monitor
Why it's wrong here
Model Monitor detects data drift and quality issues in deployed models; it has no role in adjusting instance counts. It is tempting because it is a real-time endpoint feature that improves operational reliability, correct when you need to alert on model degradation, not when you need demand-based scaling.
- ✓
Auto Scaling for SageMaker endpoints
Why this is correct
Application Auto Scaling for SageMaker endpoints adjusts the instance count of a production variant in response to CloudWatch metrics such as InvocationsPerInstance, matching capacity to variable demand. This satisfies the requirement to scale automatically while preserving performance during peaks.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.