Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

An ML engineer needs to monitor the operational health of a SageMaker endpoint, specifically the time taken for the container to process an inference request and the overhead added by SageMaker. Which two CloudWatch metrics should they examine?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ModelLatency and OverheadLatency

ModelLatency is the time taken by the model to respond, and OverheadLatency is the additional time added by SageMaker infrastructure. Invocations is count, not duration; Latency is total latency (ModelLatency + OverheadLatency).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    ModelLatency and 4XXError

    Why it's wrong here

    ModelLatency captures container processing time but not SageMaker's own overhead, and 4XXError counts client errors rather than latency. The pair is tempting because 4XXError is a genuine endpoint health metric, but it answers error-rate monitoring, not the invocation overhead the stem asks about.

  • ✗

    Latency and 5XXError

    Why it's wrong here

    Latency gives total end-to-end response time and 5XXError counts server-side failures, so neither isolates container processing time from SageMaker overhead. It is tempting because 5XXError genuinely signals endpoint faults, which would be correct for diagnosing availability problems rather than decomposing inference latency.

  • ✓

    ModelLatency and OverheadLatency

    Why this is correct

    ModelLatency captures the time the container spends processing the inference request, while OverheadLatency measures the additional time SageMaker adds for request routing and response handling. Together they decompose total endpoint latency into model and platform components.

  • ✗

    Invocations and Latency

    Why it's wrong here

    Invocations counts requests and Latency measures end-to-end time, but neither separates container processing time from SageMaker-added overhead. It is tempting because Latency is genuinely an endpoint timing metric, which would be correct for overall responsiveness rather than the ModelLatency and OverheadLatency split the question demands.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.