Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

A team monitors a production endpoint and notices a sudden increase in 5XXError count. Which of the following is the most likely cause?

⚠ Common exam trap

In AWS, 5XX errors on a SageMaker endpoint indicate server-side failures (e.g., container crash, out-of-memory). A common trap is to confuse 5XX errors with client-side errors like throttling (HTTP 429) or input format issues (HTTP 400), but only 5XX errors point to a problem within the model container or inference code.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The model container is out of memory or crashing

A sudden increase in 5XX errors, particularly HTTP 503 or 502, typically indicates that the model container is failing to process requests due to resource exhaustion (e.g., OOM kills) or a crash in the inference process. In a production ML endpoint, such errors often stem from the container running out of memory, leading to the container being terminated by the orchestrator (e.g., Kubernetes OOMKill) or the application crashing internally, which directly causes 5XX responses.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The endpoint is under-provisioned and requests are throttled

    Why it's wrong here

    Throttling produces 429 Too Many Requests responses, not 5XXError, because the request is rejected before reaching the model. Provisioning capacity is the right remedy when 429 throttling counts rise, which is a distinct CloudWatch metric from 5XXError.

  • ✗

    The input data format has changed

    Why it's wrong here

    A changed input format typically yields 4XX client errors or successful predictions, not server-side 5XX faults, unless it triggers an unhandled exception. Schema validation is the right control when diagnosing malformed payloads, but it does not explain a 5XXError spike.

  • ✓

    The model container is out of memory or crashing

    Why this is correct

    5XX errors are server-side failures returned by the endpoint itself, so the fault lies in the container rather than the client request. Memory exhaustion or a container crash prevents the model server from completing inference, producing exactly this error class, whereas throttling or bad input would surface as 4XX responses.

  • ✗

    The model is returning predictions with high latency

    Why it's wrong here

    High latency alone does not raise 5XXError; it surfaces as elevated model latency or timeout metrics. Latency tuning is the correct focus when response times breach service-level objectives, not when the endpoint reports server-side failures.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.