MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security
A company is deploying a SageMaker real-time endpoint and needs to monitor inference latency. Which THREE metrics are available from SageMaker for this purpose? (Choose THREE.)
⚠ Common exam trap
Watch out — candidates often confuse Invocations (a request count metric) or MemoryUtilization (a resource utilization metric) with latency metrics, but SageMaker specifically provides three distinct latency-focused metrics: Latency, ModelLatency, and OverheadLatency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
OverheadLatency
Option A, OverheadLatency, is correct because SageMaker publishes it as the time spent on overhead outside the model itself (for example, request routing and response processing) in the endpoint's invocation path, so it directly contributes to observed inference latency. Option C, ModelLatency, is correct because it measures the interval the container spends processing the request, which is the core inference latency component reported by SageMaker. Option E, Latency, is correct because SageMaker reports the total end-to-end time from when the request is received to when the response is returned, which is the primary inference latency metric for a real-time endpoint. Option B, Invocations, is not a latency metric; it counts the number of requests sent to the endpoint. Option D, MemoryUtilization, is a resource-utilization metric for the instance/container, not an inference latency measurement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
OverheadLatency
Why this is correct
OverheadLatency measures the time SageMaker spends on request handling outside model inference, such as preprocessing and network overhead. It directly satisfies the stem's inference-latency monitoring requirement, since total latency comprises model latency plus overhead, letting teams isolate endpoint-side delays from model execution time.
- ✗
Invocations
Why it's wrong here
Invocations counts requests reaching the endpoint, not the time taken to respond, so it cannot measure inference latency. It is genuinely useful for throughput and error-rate dashboards, and would be the right metric if the question asked about request volume rather than latency.
- ✓
ModelLatency
Why this is correct
ModelLatency measures the interval between when the model container receives a request and when it returns the response, excluding network and overhead time. This directly satisfies the requirement to monitor inference latency on a real-time endpoint, isolating the model's own processing duration from other stages of the request path.
- ✗
MemoryUtilization
Why it's wrong here
MemoryUtilization reports container memory consumption on the instance, not the interval between request arrival and response, so it cannot quantify inference latency. It belongs in capacity and out-of-memory monitoring, and would be correct if the stem asked about resource saturation rather than response timing.
- ✓
Latency
Why this is correct
Latency measures the time between a request arriving at the endpoint and the response being returned, directly satisfying the inference-latency monitoring requirement. SageMaker publishes it automatically to Amazon CloudWatch for real-time endpoints, alongside ModelLatency and OverheadLatency, giving granular visibility into total versus in-model processing time.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.