Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

A company wants to reduce costs for a SageMaker real-time endpoint that has variable traffic. Which feature allows the endpoint to automatically adjust instance count based on demand?

⚠ Common exam trap

MLA-C01 often tests the confusion between cost-optimization features — candidates pick Savings Plans (a billing discount) when the question is actually about dynamic capacity adjustment via autoscaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Auto Scaling for SageMaker endpoints

Auto Scaling for SageMaker endpoints is the native capability that dynamically adjusts the number of instances behind a real-time endpoint based on CloudWatch metrics such as InvocationsPerInstance or ModelLatency. It uses Application Auto Scaling policies (target tracking or step scaling) to add instances during traffic spikes and remove them during lulls, directly reducing cost for variable workloads. Savings Plans and Inference Recommender do not perform runtime scaling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    SageMaker Savings Plans

    Why it's wrong here

    Savings Plans are a billing commitment giving discounted rates on compute usage; they do not change instance count in response to traffic. It is tempting because they genuinely reduce cost for predictable, steady workloads, but the scenario requires automatic scaling, which Savings Plans cannot perform.

  • ✗

    SageMaker Inference Recommender

    Why it's wrong here

    Inference Recommender analyses workloads and suggests instance types and configurations; it does not scale a live endpoint's instance count. It is tempting because it optimises instance selection for cost and performance, which is correct during initial deployment planning, not for automatic runtime scaling under variable demand.

  • ✗

    SageMaker Model Monitor

    Why it's wrong here

    Model Monitor detects data drift and quality issues in deployed models; it has no role in adjusting instance counts. It is tempting because it is a real-time endpoint feature that improves operational reliability, correct when you need to alert on model degradation, not when you need demand-based scaling.

  • ✓

    Auto Scaling for SageMaker endpoints

    Why this is correct

    Application Auto Scaling for SageMaker endpoints adjusts the instance count of a production variant in response to CloudWatch metrics such as InvocationsPerInstance, matching capacity to variable demand. This satisfies the requirement to scale automatically while preserving performance during peaks.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.