mediumMultiple Choice
MLA-C01 Practice Question: A company deploys a real-time inference endpoint…
A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?
⚠ Common exam trap
Many exam-takers confuse scaling the SageMaker endpoint with scaling the instance size, thinking a larger instance (Option B) is the simplest solution, but the AWS exam tests understanding of dynamic, cost-optimized scaling using target tracking policies based on CloudWatch metrics rather than static over-provisioning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.
SageMaker endpoints support automatic scaling through target tracking scaling policies based on Amazon CloudWatch metrics like InvocationsPerInstance. This allows the endpoint to dynamically adjust the number of instances in response to real-time traffic spikes, scaling out when demand increases and scaling in when it decreases, which optimizes cost by only paying for the capacity needed at any given time.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to batch transform for all inference requests.
Why it's wrong here
Batch transform processes stored data on a schedule, returning predictions only after the whole job finishes, so a customer-facing application receives no synchronous response to each request. It suits offline scoring of accumulated records, not interactive inference.
- ✗
Use a larger instance type to handle peak traffic.
Why it's wrong here
A larger instance type raises the ceiling for one endpoint but remains statically provisioned, so unpredictable spikes beyond that capacity still fail while quiet periods waste spend. It suits steady, predictable workloads where peak demand is known in advance.
- ✓
Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.
Why this is correct
Target tracking scaling adjusts instance count automatically against a CloudWatch metric such as InvocationsPerInstance, matching capacity to unpredictable spikes without manual intervention. It satisfies the stem's dual constraint: automatic scaling under variable load while minimising cost, since capacity shrinks during quiet periods rather than provisioning for peak.
- ✗
Deploy multiple models behind an Application Load Balancer.
Why it's wrong here
An Application Load Balancer distributes requests across fixed endpoints but performs no automatic scaling of SageMaker instance count, so spikes still overwhelm provisioned capacity and idle periods still bill. It suits routing across already-provisioned heterogeneous services.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.