Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

An ML team uses SageMaker to deploy a model for real-time inference. They want to monitor and improve cost efficiency. Which THREE actions should they take? (Select THREE.)

⚠ Common exam trap

It's easy for candidates to confuse monitoring (Option C) with cost optimization, or they mistakenly apply Spot Training (Option D) to inference endpoints, not realizing that Spot instances are only supported for training and not for real-time inference due to interruption risk.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker Inference Recommender to find the optimal instance type and count

Option A is correct because SageMaker Inference Recommender runs load tests and benchmarks to recommend the optimal instance type and instance count for a real-time endpoint, directly improving cost efficiency by avoiding over-provisioning. Option B is correct because configuring auto-scaling on a SageMaker endpoint adjusts the number of instances to match traffic demand, so the team pays only for capacity actually needed during peaks and troughs. Option E is correct because SageMaker Savings Plans offer discounted pricing (up to 64% off) in exchange for a committed hourly spend, reducing the cost of steady-state real-time inference workloads. Option C is not a cost-efficiency action; a CloudWatch dashboard for latency is an observability tool and does not by itself reduce spend. Option D is incorrect because Managed Spot Training applies to training jobs, not to real-time inference endpoint instances, which cannot use spot capacity for persistent endpoints.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use SageMaker Inference Recommender to find the optimal instance type and count

    Why this is correct

    Inference Recommender benchmarks candidate instance types and counts against the model's actual traffic and latency requirements, recommending the most cost-effective configuration. This directly addresses right-sizing, which is the primary lever for reducing real-time inference cost before scaling or commitment discounts.

  • ✓

    Enable auto-scaling to adjust the number of instances based on demand

    Why this is correct

    Auto-scaling adjusts instance count dynamically to match demand, so the endpoint provisions capacity only when needed and scales in during idle periods. This directly satisfies the cost-efficiency goal by eliminating over-provisioned instances during low-traffic windows.

  • ✗

    Create a CloudWatch dashboard to monitor endpoint latency

    Why it's wrong here

    A latency dashboard surfaces performance, not cost; it identifies no spend driver and triggers no efficiency action. It is tempting because CloudWatch also exposes invocation and instance metrics that inform capacity decisions. It would be correct when diagnosing endpoint response-time or availability problems rather than improving cost efficiency.

  • ✗

    Use SageMaker Managed Spot Training for endpoint instances

    Why it's wrong here

    Managed Spot Training applies to training jobs, not endpoint instances; real-time endpoints run on persistent instances and cannot use spot capacity. It is tempting because spot capacity genuinely cuts compute cost, but only for interruptible training workloads. It would be correct when reducing the cost of model training rather than inference hosting.

  • ✓

    Purchase SageMaker Savings Plans for a discounted rate

    Why this is correct

    Savings Plans commit to a consistent hourly SageMaker spend in exchange for discounted rates on eligible usage, including real-time inference instances. This reduces the per-hour cost of steady baseline capacity, complementing right-sizing and auto-scaling for the cost-efficiency objective.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.