MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security
An ML engineer needs to monitor the operational health of a SageMaker endpoint, specifically the time taken for the container to process an inference request and the overhead added by SageMaker. Which two CloudWatch metrics should they examine?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ModelLatency and OverheadLatency
ModelLatency is the time taken by the model to respond, and OverheadLatency is the additional time added by SageMaker infrastructure. Invocations is count, not duration; Latency is total latency (ModelLatency + OverheadLatency).
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
ModelLatency and 4XXError
Why it's wrong here
ModelLatency captures container processing time but not SageMaker's own overhead, and 4XXError counts client errors rather than latency. The pair is tempting because 4XXError is a genuine endpoint health metric, but it answers error-rate monitoring, not the invocation overhead the stem asks about.
- ✗
Latency and 5XXError
Why it's wrong here
Latency gives total end-to-end response time and 5XXError counts server-side failures, so neither isolates container processing time from SageMaker overhead. It is tempting because 5XXError genuinely signals endpoint faults, which would be correct for diagnosing availability problems rather than decomposing inference latency.
- ✓
ModelLatency and OverheadLatency
Why this is correct
ModelLatency captures the time the container spends processing the inference request, while OverheadLatency measures the additional time SageMaker adds for request routing and response handling. Together they decompose total endpoint latency into model and platform components.
- ✗
Invocations and Latency
Why it's wrong here
Invocations counts requests and Latency measures end-to-end time, but neither separates container processing time from SageMaker-added overhead. It is tempting because Latency is genuinely an endpoint timing metric, which would be correct for overall responsiveness rather than the ModelLatency and OverheadLatency split the question demands.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.