Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

A company uses SageMaker Inference Recommender to select the optimal endpoint configuration. After running the recommender, they receive a recommendation for a specific instance type and initial instance count. What should they do next to optimize costs over time?

⚠ Common exam trap

The trap is treating the Inference Recommender output as a final, optimal configuration — candidates forget it is a starting point and that ongoing cost optimization requires auto-scaling to match actual traffic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set up auto-scaling with a target tracking policy based on the recommended metric

Inference Recommender provides an initial right-sized instance type and count based on load testing, but traffic varies over time. Setting up auto-scaling with a target tracking policy (e.g., on InvocationsPerInstance or CPU utilization) lets the endpoint scale in during low traffic and out during peaks, optimizing cost continuously rather than paying for a static over-provisioned fleet.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the recommended configuration without changes, as it is already optimal

    Why it's wrong here

    Inference Recommender produces a starting instance type and count from a load test, not a continuously optimal configuration; traffic patterns shift, so the endpoint still needs ongoing monitoring and adjustment. It is tempting because the recommendation is evidence-based, but it suits initial deployment sizing rather than long-term cost optimisation.

  • ✗

    Purchase a Savings Plan for the recommended instance type to reduce hourly cost

    Why it's wrong here

    A Savings Plan discounts compute spend but does not change the endpoint's instance type or count, so it cannot address over-provisioning as demand changes. It is tempting because it lowers hourly rates, and it would be correct once the endpoint's steady-state capacity is known and stable.

  • ✓

    Set up auto-scaling with a target tracking policy based on the recommended metric

    Why this is correct

    Inference Recommender only supplies a static instance type and count, so costs stay fixed regardless of traffic. A target tracking auto-scaling policy adjusts instance count dynamically against the recommended metric, matching capacity to actual load and eliminating idle spend over time.

  • ✗

    Manually adjust the instance count daily based on observed traffic

    Why it's wrong here

    Manual daily adjustment cannot track intraday traffic variation and adds operational overhead without automatic scaling. It is tempting because it gives direct control over instance count, and it would suit a workload with predictable, infrequent changes rather than continuous demand fluctuation.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.