Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

A company is deploying a SageMaker real-time endpoint and needs to monitor inference latency. Which THREE metrics are available from SageMaker for this purpose? (Choose THREE.)

⚠ Common exam trap

Watch out — candidates often confuse Invocations (a request count metric) or MemoryUtilization (a resource utilization metric) with latency metrics, but SageMaker specifically provides three distinct latency-focused metrics: Latency, ModelLatency, and OverheadLatency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

OverheadLatency

Option A, OverheadLatency, is correct because SageMaker publishes it as the time spent on overhead outside the model itself (for example, request routing and response processing) in the endpoint's invocation path, so it directly contributes to observed inference latency. Option C, ModelLatency, is correct because it measures the interval the container spends processing the request, which is the core inference latency component reported by SageMaker. Option E, Latency, is correct because SageMaker reports the total end-to-end time from when the request is received to when the response is returned, which is the primary inference latency metric for a real-time endpoint. Option B, Invocations, is not a latency metric; it counts the number of requests sent to the endpoint. Option D, MemoryUtilization, is a resource-utilization metric for the instance/container, not an inference latency measurement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    OverheadLatency

    Why this is correct

    OverheadLatency measures the time SageMaker spends on request handling outside model inference, such as preprocessing and network overhead. It directly satisfies the stem's inference-latency monitoring requirement, since total latency comprises model latency plus overhead, letting teams isolate endpoint-side delays from model execution time.

  • ✗

    Invocations

    Why it's wrong here

    Invocations counts requests reaching the endpoint, not the time taken to respond, so it cannot measure inference latency. It is genuinely useful for throughput and error-rate dashboards, and would be the right metric if the question asked about request volume rather than latency.

  • ✓

    ModelLatency

    Why this is correct

    ModelLatency measures the interval between when the model container receives a request and when it returns the response, excluding network and overhead time. This directly satisfies the requirement to monitor inference latency on a real-time endpoint, isolating the model's own processing duration from other stages of the request path.

  • ✗

    MemoryUtilization

    Why it's wrong here

    MemoryUtilization reports container memory consumption on the instance, not the interval between request arrival and response, so it cannot quantify inference latency. It belongs in capacity and out-of-memory monitoring, and would be correct if the stem asked about resource saturation rather than response timing.

  • ✓

    Latency

    Why this is correct

    Latency measures the time between a request arriving at the endpoint and the response being returned, directly satisfying the inference-latency monitoring requirement. SageMaker publishes it automatically to Amazon CloudWatch for real-time endpoints, alongside ModelLatency and OverheadLatency, giving granular visibility into total versus in-model processing time.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.