Courseiva
mediumMultiple Select

MLA-C01 Practice Question: An ML engineer is setting up monitoring for a…

An ML engineer is setting up monitoring for a SageMaker endpoint. Which THREE metrics should be monitored to detect performance issues? (Select THREE.)

⚠ Common exam trap

A common mix-up: candidates confuse throughput metrics (like invocations per second) with performance health indicators, but the question specifically asks for metrics that detect performance issues, not just operational statistics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Model latency

Model latency (A) is correct because it directly measures the time the endpoint takes to respond to inference requests, and rising latency is a primary indicator of performance degradation. CPUUtilization (C) is correct because high CPU usage on the hosting instances can throttle inference processing and cause slower responses or timeouts. MemoryUtilization (D) is correct because excessive memory consumption can lead to swapping, out-of-memory errors, or container restarts that degrade endpoint performance. Invocations per second (B) reflects traffic volume rather than endpoint health, so it does not by itself indicate a performance issue. DiskWriteBytes (E) is a storage I/O metric that is largely irrelevant to the in-memory inference workload of a SageMaker endpoint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Model latency

    Why this is correct

    Model latency measures the time the endpoint takes to return a prediction, so rising values directly reveal inference performance degradation. It satisfies the requirement to detect endpoint performance issues, exposing slowdowns before they breach application response-time expectations.

  • ✗

    Invocations per second

    Why it's wrong here

    Invocations per second is a throughput volume metric, not a performance-failure indicator; ModelLatency and Invocation4XXErrors reveal degradation. It tempts because invocation counts appear in endpoint dashboards, and it would be correct for capacity planning or autoscaling triggers rather than detecting latency or error spikes.

  • ✓

    CPUUtilization

    Why this is correct

    CPUUtilization tracks the percentage of CPU instances consumed by the endpoint, exposing compute saturation that degrades inference latency. It satisfies the stem's performance-monitoring requirement by revealing resource exhaustion before requests time out, complementing ModelLatency and Invocations metrics for a complete operational view.

  • ✓

    MemoryUtilization

    Why this is correct

    MemoryUtilization tracks the percentage of memory consumed by the endpoint instance, exposing memory exhaustion that degrades inference throughput or triggers out-of-memory failures. Monitoring it alongside CPU and GPU utilisation satisfies the stem's requirement to detect performance issues, since sustained high memory usage signals the instance is undersized for the deployed model's payload.

  • ✗

    DiskWriteBytes

    Why it's wrong here

    DiskWriteBytes measures instance-level disk throughput, not endpoint inference behaviour, so it cannot reveal model latency or error problems. It tempts because it is a real CloudWatch metric, and it would be relevant when monitoring the storage performance of the underlying instance rather than the endpoint itself.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.