Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A healthcare company uses Amazon SageMaker to…

A healthcare company uses Amazon SageMaker to deploy a real-time inference endpoint for a diagnostic model. The endpoint is configured with a single ml.p3.2xlarge instance. The model processes patient data and returns a risk score. Recently, the endpoint has been experiencing intermittent 504 errors along with increased latency. The team uses Amazon CloudWatch to monitor the endpoint's InvocationsPerInstance and ModelLatency metrics. They observe that InvocationsPerInstance is well below the throttling threshold, but ModelLatency shows periodic spikes lasting 5-10 seconds. The endpoint's CPU utilization remains below 60%, but memory utilization occasionally spikes to 90% during those spikes. The team has checked the inference code and found no obvious memory leaks or performance bottlenecks in the custom logic. The model itself is a deep neural network hosted using Apache MXNet. The team suspects that the issue might be related to resource contention or an external dependency. What should the team do FIRST to diagnose and resolve the issue?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the instance type to a more memory-intensive instance like ml.p3.8xlarge to handle memory spikes.

The endpoint memory spikes to 90% and correlates with 504 errors and latency spikes, so the team should first resolve the resource pressure by moving to a larger memory instance such as ml.p3.8xlarge. The current correct option D is technically inaccurate: Amazon SageMaker Debugger rules and profiling are for training jobs, not real-time inference endpoints. Option A does not address memory pressure. Option C is for data drift and model quality, not performance diagnostics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Implement request batching to increase throughput and reduce the number of inference requests.

    Why it's wrong here

    Request batching improves throughput but does not diagnose memory utilization and can increase memory usage.

  • ✓

    Increase the instance type to a more memory-intensive instance like ml.p3.8xlarge to handle memory spikes.

    Why this is correct

    Increasing to a memory-intensive instance directly addresses the observed memory spikes and associated endpoint timeouts.

  • ✗

    Set up SageMaker Model Monitor to track data drift and model quality metrics.

    Why it's wrong here

    SageMaker Model Monitor is for data drift and model quality monitoring, not real-time performance debugging.

  • ✗

    Enable SageMaker Debugger rules and profiling to monitor memory and CPU utilization at a fine-grained level during inference.

    Why it's wrong here

    SageMaker Debugger profiling is not available for inference endpoints; it is a training-job debugging and profiling tool.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.