Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

A company wants to reduce costs for a production SageMaker endpoint that has predictable traffic patterns. They have purchased a Savings Plan. What additional step can they take to further optimize costs while maintaining performance?

⚠ Common exam trap

MLA-C01 often tests the misconception that a Savings Plan alone fully optimizes cost — candidates forget that right-sizing the instance fleet via Inference Recommender is the complementary step that preserves performance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker Inference Recommender to right-size the endpoint

SageMaker Inference Recommender runs load tests against candidate instance types and configurations to identify the most cost-effective endpoint that still meets latency and throughput requirements. Since the Savings Plan already discounts compute, right-sizing the underlying instance fleet is the remaining lever for cost optimization without sacrificing performance. This directly addresses the 'maintaining performance' constraint that rules out simply shrinking capacity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use SageMaker Inference Recommender to right-size the endpoint

    Why this is correct

    Inference Recommender profiles the model against candidate instance types and configurations, identifying the cheapest instance that still meets latency and throughput targets. This right-sizes the endpoint, so the Savings Plan discount applies to a smaller, better-matched instance, compounding the cost reduction.

  • ✗

    Reduce the number of instances to one, regardless of load

    Why it's wrong here

    Collapsing to a single instance caps total throughput and concurrent invocations, so the endpoint cannot serve peak load even when traffic is predictable. It is tempting because fewer instances directly cut hourly charges, yet the correct approach scales instance count to the forecast curve, which is what auto-scaling with scheduled policies is designed for.

  • ✗

    Switch from real-time to batch inference

    Why it's wrong here

    Batch Transform cannot serve synchronous, low-latency requests, so any application expecting real-time responses breaks. It is tempting because batch jobs avoid continuously billed endpoint hours, yet batch inference suits offline scoring of stored datasets, whereas a production endpoint must answer requests immediately.

  • ✗

    Disable auto-scaling

    Why it's wrong here

    Disabling auto-scaling removes the endpoint's ability to add instances during demand spikes, so throughput degrades whenever traffic exceeds the fixed capacity. It is tempting because scaling activity itself incurs no direct charge, yet auto-scaling exists precisely to match provisioned capacity to load, and predictable patterns still need scheduled scaling to preserve performance.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.