easyMultiple Choice
MLA-C01 Practice Question: A company uses Amazon SageMaker to deploy a…
A company uses Amazon SageMaker to deploy a real-time inference endpoint. They notice increased latency in predictions during peak hours. Which should they investigate first to address the issue?
⚠ Common exam trap
Many exam-takers confuse training infrastructure (instance type, artifact size) with inference infrastructure, or assume that data labeling quality affects inference speed, when the immediate cause of peak-hour latency is almost always insufficient endpoint capacity due to misconfigured auto-scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Review the endpoint auto-scaling policy
Increased latency during peak hours is a classic symptom of insufficient compute capacity to handle the request volume. The first step is to review the endpoint's auto-scaling policy to ensure it is configured to scale out instances proactively or reactively based on a relevant metric like 'SageMakerVariantInvocationsPerInstance'. If the policy has a high cooldown period or a low target metric value, it may not add instances quickly enough, causing requests to queue and latency to spike.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Review the endpoint auto-scaling policy
Why this is correct
Peak-hour latency typically indicates insufficient capacity, so reviewing the endpoint auto-scaling policy reveals whether instance counts and target-tracking thresholds keep pace with demand. This satisfies the requirement to investigate the most likely cause first, before examining model artefacts or payload sizes.
- ✗
Check the data labeling job status
Why it's wrong here
Data labelling job status concerns preparing training datasets and has no bearing on the runtime latency of a deployed inference endpoint. It is tempting because labelling is part of the SageMaker workflow, but it would be the correct area to check when investigating training data quality or a stalled ground-truth job, not endpoint performance.
- ✗
Modify the training instance type
Why it's wrong here
Training instance type affects model fitting cost and duration, not the latency of an already-deployed real-time endpoint during peak traffic. It is tempting because instance sizing influences performance generally, but the correct choice here is endpoint-side scaling or instance selection, which directly governs inference response times under load.
- ✗
Increase the model artifact size
Why it's wrong here
A larger model artifact increases download and load time and typically raises per-inference latency, so enlarging it worsens the symptom rather than addressing peak-hour latency. Artifact size matters when optimising cold-start or container image pull times, not for handling concurrent real-time inference traffic.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.