Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A company deploys a real-time inference endpoint…

A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?

⚠ Common exam trap

Many exam-takers confuse scaling the SageMaker endpoint with scaling the instance size, thinking a larger instance (Option B) is the simplest solution, but the AWS exam tests understanding of dynamic, cost-optimized scaling using target tracking policies based on CloudWatch metrics rather than static over-provisioning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.

SageMaker endpoints support automatic scaling through target tracking scaling policies based on Amazon CloudWatch metrics like InvocationsPerInstance. This allows the endpoint to dynamically adjust the number of instances in response to real-time traffic spikes, scaling out when demand increases and scaling in when it decreases, which optimizes cost by only paying for the capacity needed at any given time.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch to batch transform for all inference requests.

    Why it's wrong here

    Batch transform processes stored data on a schedule, returning predictions only after the whole job finishes, so a customer-facing application receives no synchronous response to each request. It suits offline scoring of accumulated records, not interactive inference.

  • ✗

    Use a larger instance type to handle peak traffic.

    Why it's wrong here

    A larger instance type raises the ceiling for one endpoint but remains statically provisioned, so unpredictable spikes beyond that capacity still fail while quiet periods waste spend. It suits steady, predictable workloads where peak demand is known in advance.

  • ✓

    Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.

    Why this is correct

    Target tracking scaling adjusts instance count automatically against a CloudWatch metric such as InvocationsPerInstance, matching capacity to unpredictable spikes without manual intervention. It satisfies the stem's dual constraint: automatic scaling under variable load while minimising cost, since capacity shrinks during quiet periods rather than provisioning for peak.

  • ✗

    Deploy multiple models behind an Application Load Balancer.

    Why it's wrong here

    An Application Load Balancer distributes requests across fixed endpoints but performs no automatic scaling of SageMaker instance count, so spikes still overwhelm provisioned capacity and idle periods still bill. It suits routing across already-provisioned heterogeneous services.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.