mediumMultiple Choice
MLA-C01 Practice Question: After deploying a model to a SageMaker endpoint,…
After deploying a model to a SageMaker endpoint, the operations team notices high inference latency. They suspect it is due to insufficient instance capacity. Which first step should they take to diagnose the issue?
⚠ Common exam trap
It's easy for candidates to confuse SageMaker Debugger (for training debugging) with inference monitoring tools, or they assume CloudTrail can provide performance metrics, when in fact CloudWatch is the correct service for real-time endpoint health and capacity diagnostics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Review Amazon CloudWatch metrics for the endpoint, such as CPUUtilization and Invocations.
Amazon CloudWatch metrics for a SageMaker endpoint, such as `CPUUtilization`, `MemoryUtilization`, and `Invocations`, directly indicate whether the instance is overloaded. High `CPUUtilization` combined with a high `Invocations` count and increased latency strongly suggests insufficient instance capacity. This is the standard first diagnostic step for capacity-related performance issues.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Check AWS CloudTrail logs for API errors.
Why it's wrong here
CloudTrail records control-plane API calls such as endpoint configuration changes, not per-invocation latency or instance saturation. It is tempting because CloudTrail is the go-to audit service for API activity, but diagnosing capacity-driven latency requires CloudWatch metrics like CPUUtilization, memory and ModelLatency on the endpoint itself.
- ✗
Use Amazon SageMaker Debugger to analyze inference performance.
Why it's wrong here
Debugger analyses training jobs — tensors, gradients and resource utilisation during fit — not live endpoint inference performance. It is tempting because Debugger surfaces resource bottlenecks, but that applies to training runs; endpoint latency diagnosis needs CloudWatch invocation metrics such as ModelLatency and CPUUtilization to confirm capacity shortfall.
- ✓
Review Amazon CloudWatch metrics for the endpoint, such as CPUUtilization and Invocations.
Why this is correct
CloudWatch publishes per-endpoint CPUUtilization and Invocations, letting the team confirm whether the instance is saturated before changing anything. These metrics directly test the capacity hypothesis, distinguishing genuine under-provisioning from model or payload latency without redeploying.
- ✗
Retrain the model with more training data.
Why it's wrong here
Retraining with more data alters model quality, not serving capacity, so latency caused by insufficient instance resources persists. It is tempting because retraining is a familiar remedy for poor model outcomes, but it addresses accuracy, whereas the stem's symptom is throughput and resource saturation, diagnosed via CloudWatch endpoint metrics.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.