mediumMultiple Choice
MLA-C01 Practice Question: A media company uses SageMaker endpoints to serve…
A media company uses SageMaker endpoints to serve a model that predicts video engagement. They have two production variants: Variant A (ml.c5.large) for regular traffic and Variant B (ml.c5.xlarge) for burst traffic. They use weighted routing (90% to A, 10% to B). Recently, during peak hours, Variant A's latency increase causes many requests to time out. The metrics show that both variants are under similar CPU load, but the number of concurrent requests to Variant A is very high. The team wants to ensure that burst traffic is handled properly without manual intervention. What should they do?
⚠ Common exam trap
MLA-C01 often tests the difference between reactive alarm-based scaling and proactive target tracking — candidates pick CloudWatch alarms because they sound more controlled, missing that target tracking is the managed, automatic solution for concurrency-based scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure Application Auto Scaling for each variant with a target tracking scaling policy based on the number of concurrent requests per instance.
Application Auto Scaling with a target tracking policy based on concurrent requests per instance automatically adds instances when concurrency rises, directly addressing the high concurrent request load on Variant A. This is the managed, no-manual-intervention solution that scales each variant independently based on the actual bottleneck metric.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the traffic weight to Variant B to 70% and reduce Variant A to 30%.
Why it's wrong here
Reweighting to Variant B shifts traffic but leaves Variant A's concurrency bottleneck unaddressed and is a manual change, not automatic. It tempts because weighted routing between variants is the standard mechanism for splitting traffic when variant capacities are genuinely mismatched.
- ✓
Configure Application Auto Scaling for each variant with a target tracking scaling policy based on the number of concurrent requests per instance.
Why this is correct
Concurrent-request target tracking scales each variant independently on the actual bottleneck, since CPU load is similar but Variant A's request concurrency is high. This removes manual intervention and lets Variant B absorb burst traffic as routing shifts.
- ✗
Set a CloudWatch alarm on Variant A's p99 latency and trigger a step scaling policy to add instances.
Why it's wrong here
A p99 latency alarm with step scaling reacts after latency has already risen, and instance provisioning still takes minutes, so timeouts continue. It tempts because CloudWatch-driven step scaling is the correct pattern when metrics lead demand and capacity changes are fast.
- ✗
Create a separate endpoint for burst traffic and route peak traffic to it via DNS.
Why it's wrong here
A separate endpoint with DNS routing is manual and static, so it cannot respond automatically to peak-hour bursts. It tempts because dedicated endpoints with DNS failover suit scenarios where traffic separation is planned and predictable rather than dynamic.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.