Courseiva
Question 484 of 835
hardMultiple ChoiceObjective-mapped

MLA-C01 Practice Question: A team deploys a machine learning model using a…

A team deploys a machine learning model using a SageMaker endpoint with an ML.T4 instance. After a week, they notice that the endpoint's CPU utilization is consistently below 10% and latency is low. However, the endpoint is incurring high costs. Which action should the team take to reduce costs while maintaining the ability to serve traffic?

⚠ Common exam trap

It's easy for candidates to assume reducing instance count or switching endpoint types (multi-model, async) will lower costs, but they overlook that provisioned instances always incur hourly charges, whereas serverless charges only for actual compute usage, making it the optimal choice for consistently low-utilization endpoints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Migrate to a SageMaker Serverless Inference endpoint

The endpoint's CPU utilization is consistently below 10% with low latency, indicating that traffic is sparse and the instance is severely underutilized. SageMaker Serverless Inference automatically scales compute resources based on request volume and charges only for the compute time consumed per inference, eliminating idle costs. This makes it the most cost-effective choice for low-utilization workloads while still maintaining the ability to serve traffic on demand.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Switch to a multi-model endpoint to share instances across models

    Why it's wrong here

    Multi-model endpoints still have fixed instances; cost savings are limited.

  • Reduce the number of instances to one

    Why it's wrong here

    Reducing instances may cause high latency during traffic spikes.

  • Migrate to a SageMaker Serverless Inference endpoint

    Why this is correct

    Serverless endpoints scale to zero when idle, reducing cost.

  • Implement an asynchronous inference endpoint

    Why it's wrong here

    Asynchronous inference is for batch processing, not real-time.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Last reviewed: Jul 4, 2026

Question Discussion

Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.

Loading comments…

Sign in to join the discussion.

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.