Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A company uses SageMaker endpoints with…

A company uses SageMaker endpoints with auto-scaling based on CPU utilization. During a flash sale, latency increases despite low CPU. What should be done?

⚠ Common exam trap

Many candidates assume CPU utilization is always the best scaling metric for compute-bound workloads, but the MLA-C01 exam specifically tests the understanding that inference endpoints can be I/O-bound, making request count or memory utilization more appropriate for auto-scaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a custom metric such as memory utilization or request count for auto-scaling

CPU utilization is a poor scaling metric for inference workloads that are I/O or memory-bound. During a flash sale, increased request concurrency can cause queuing and latency spikes even when CPU is low. Using a custom metric like request count per instance or memory utilization directly reflects the load on the inference endpoint, enabling the Application Auto Scaling target tracking policy to scale out proactively before latency degrades.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a custom metric such as memory utilization or request count for auto-scaling

    Why this is correct

    CPU utilisation stays low while latency rises, so the scaling policy tracks a metric unrelated to the actual bottleneck. Scaling on memory utilisation or request count instead triggers additional instances when the real constraint saturates, restoring latency during the flash sale.

  • ✗

    Increase the instance size

    Why it's wrong here

    Increasing instance size raises per-instance memory and compute, yet the stem shows CPU is low, so the bottleneck is concurrent request handling, not instance capacity. Larger instances suit steady high-throughput workloads, whereas the flash sale needs the endpoint to scale out on a concurrency-based metric.

  • ✗

    Disable auto-scaling and use a larger instance

    Why it's wrong here

    Disabling auto-scaling removes the endpoint's ability to add instances when concurrency rises, so the flash-sale load still queues against one endpoint. A larger instance is tempting because it raises per-instance throughput, but it cannot match the burst concurrency that scaling on a concurrency metric would absorb.

  • ✗

    Switch to GPU instances

    Why it's wrong here

    GPU instances accelerate compute-bound inference; here CPU is low while latency rises, so the bottleneck lies elsewhere, such as concurrency or memory. GPUs suit models whose inference is genuinely compute-intensive, not endpoints already under-utilising CPU.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.