MLA-C01 Deployment and Orchestration of ML Workflows Practice Question
A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker AI for a fraud detection model. The model must serve predictions with consistent latency under 50 ms and the team expects traffic to fluctuate unpredictably, with occasional bursts. The engineer wants to automatically adjust the number of instances based on actual workload while minimizing cost during idle periods. Which SageMaker AI feature should the engineer configure?
⚠ Common exam trap
Many candidates confuse a one-time right-sizing recommendation from Inference Recommender with runtime automatic scaling that responds to live traffic.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
SageMaker AI automatic scaling with a target-tracking policy based on the SageMakerVariantInvocationsPerInstance metric
Automatic scaling with target-tracking on the invocations-per-instance metric is the correct approach because it continuously adjusts the number of instances to match actual traffic, ensuring latency targets are met while scaling down during idle periods to save cost. The other options either provide static recommendations or are designed for different inference patterns.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
SageMaker AI multi-model endpoints with a shared inference container
Why it's wrong here
Multi-model endpoints host multiple models behind one endpoint to reduce hosting costs, but they do not provide automatic scaling based on traffic. The number of instances remains fixed unless separately scaled. This does not address the need to dynamically adjust capacity for a single model under variable load.
- ✓
SageMaker AI automatic scaling with a target-tracking policy based on the SageMakerVariantInvocationsPerInstance metric
Why this is correct
Target-tracking automatic scaling uses CloudWatch metrics like SageMakerVariantInvocationsPerInstance to adjust instance count dynamically, matching capacity to demand. This directly addresses unpredictable traffic and minimizes cost during low usage while maintaining latency. It is the native SageMaker AI scaling mechanism for real-time endpoints.
- ✗
SageMaker AI Inference Recommender to select the optimal instance type and count
Why it's wrong here
Inference Recommender is a one-time right-sizing tool that benchmarks models and recommends instance types; it does not continuously adjust capacity at runtime. It cannot respond to fluctuating traffic or bursts, so it fails to maintain consistent latency automatically. It is used before deployment, not for dynamic scaling.
- ✗
SageMaker AI Asynchronous Inference with an auto-scaling policy on queue depth
Why it's wrong here
Asynchronous Inference is designed for large payloads and long processing times, not sub-50 ms real-time predictions. While it supports auto scaling, it introduces queuing and is not suitable for low-latency synchronous fraud detection. The scenario requires immediate responses, so this approach is inappropriate.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.