MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security
A team receives alerts that their SageMaker endpoint latency has increased significantly. They check CloudWatch metrics and see Invocations rising, but ModelLatency remains stable. Which metric should they investigate to find the source of the increased latency?
⚠ Common exam trap
Candidates often assume increased Invocations directly cause higher ModelLatency, but the exam tests the distinction between inference time and infrastructure overhead, leading them to incorrectly select ModelLatency instead of OverheadLatency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
OverheadLatency
OverheadLatency measures the time taken by the SageMaker infrastructure to handle requests before and after model inference, including request routing, authentication, and response processing. Since ModelLatency is stable but total endpoint latency has increased, the extra time must be in the overhead component, making OverheadLatency the correct metric to investigate.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
OverheadLatency
Why this is correct
OverheadLatency measures time spent outside the model, covering request routing, queueing and response handling. Since Invocations rose while ModelLatency stayed flat, the added delay sits in this overhead, not in model execution, pinpointing the source.
- ✗
ModelLatency
Why it's wrong here
ModelLatency measures only the time the model container spends computing inferences; it is already stable, so it cannot explain the rise. It is tempting because it is the obvious latency metric, but when Invocations climb while ModelLatency holds, the delay lies in request queuing or overhead outside the model.
- ✗
5XXError
Why it's wrong here
5XXError records server-side failures returned by the endpoint, not the time requests spend waiting before inference; rising Invocations with stable ModelLatency points to overhead outside the model. It is tempting because 5XXError can reveal capacity or internal faults, but it would be correct when investigating failed responses rather than latency.
- ✗
4XXError
Why it's wrong here
4XXError counts client-side requests rejected before inference, so it cannot explain latency from rising Invocations; those requests never reach the model. It is tempting because error metrics can indicate throttling or malformed traffic, but it would be the right metric when diagnosing failed or unauthorised requests rather than queueing delay.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.