mediumMultiple Select
MLA-C01 Practice Question: A machine learning engineer is deploying a model…
A machine learning engineer is deploying a model using SageMaker and needs to ensure that the endpoint can automatically scale based on traffic patterns. Which TWO actions should the engineer take? (Choose two.)
⚠ Common exam trap
A common mix-up: candidates confuse monitoring and scaling: candidates often pick Model Monitor (Option C) because it sounds like it monitors traffic, but it is for data drift, not scaling; similarly, batch transform (Option E) is mistaken for a scaling solution when it is a separate inference mode.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Define a scaling policy using Application Auto Scaling for the SageMaker endpoint variant.
Option A is correct because SageMaker endpoint variants are scaled through Application Auto Scaling, which is the AWS service that registers the SageMaker variant as a scalable target and applies a scaling policy (target-tracking or step scaling) to adjust the desired instance count. Option B is correct because Application Auto Scaling policies are driven by Amazon CloudWatch alarms, and the InvocationsPerInstance metric is the standard SageMaker metric used to scale on traffic per instance. Option C is incorrect because SageMaker Model Monitor detects data drift and quality issues, not traffic-based scaling. Option D is incorrect because multi-model endpoints consolidate multiple models on shared infrastructure to reduce hosting cost, not to autoscale on traffic patterns. Option E is incorrect because batch transform is an offline, batch inference mechanism and does not serve a real-time, autoscaling endpoint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Define a scaling policy using Application Auto Scaling for the SageMaker endpoint variant.
Why this is correct
Application Auto Scaling registers the SageMaker endpoint variant as a scalable target and applies a target-tracking or step policy, enabling automatic replica adjustment. This satisfies the requirement for traffic-based scaling by letting the endpoint add or remove instances as load changes.
- ✓
Set up an Amazon CloudWatch alarm to trigger scaling based on the InvocationsPerInstance metric.
Why this is correct
CloudWatch publishes the InvocationsPerInstance metric for SageMaker endpoints; an alarm on it triggers the scaling policy when per-instance traffic crosses a threshold. This satisfies the automatic scaling requirement by supplying the metric-driven trigger that drives replica changes.
- ✗
Enable SageMaker Model Monitor to detect data drift.
Why it's wrong here
Model Monitor detects data and model quality drift on captured inference data; it does not adjust endpoint instance counts in response to traffic. It tempts because drift detection is essential for production monitoring, the right choice when the requirement is quality alerting rather than autoscaling.
- ✗
Configure a multi-model endpoint to serve multiple models.
Why it's wrong here
Multi-model endpoints host several models behind one endpoint to cut hosting cost; they do not themselves scale capacity with request volume. It tempts because consolidating models is correct when many low-traffic models must share infrastructure, not when traffic-driven autoscaling is the goal.
- ✗
Use SageMaker batch transform to handle variable traffic.
Why it's wrong here
Batch transform runs offline inference jobs against stored data and has no persistent endpoint, so it cannot respond to live traffic patterns. It tempts because batch transform suits large scheduled scoring workloads, which is the correct choice when real-time autoscaling is not required.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.