Courseiva
mediumMultiple Choice

MLA-C01 Using SageMaker endpoints for inference Practice Question

A company is using SageMaker endpoints for inference. To reduce costs, they want to use Automatic Scaling. However, they observe that scaling up takes several minutes, causing latency spikes during traffic bursts. What should they do to mitigate this?

⚠ Common exam trap

AWS often tests the misconception that optimizing model performance or using larger instances can eliminate scaling delays, but the real issue is the provisioning time, which requires proactive capacity management like pre-warming.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the endpoint with a target tracking scaling policy and pre-warm additional instances during expected traffic surges.

It combines a target tracking scaling policy with pre-warming additional instances during expected traffic surges. Pre-warming (or provisioning) instances ahead of time ensures that capacity is available when the burst occurs, mitigating the latency spike caused by the several-minute scaling-up delay. This approach directly addresses the cold-start problem in SageMaker auto-scaling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Optimize the model to reduce inference time.

    Why it's wrong here

    Reducing inference time shortens per-request duration but does not address the minutes-long delay to provision new endpoint instances during bursts, so latency spikes persist. It tempts because model optimisation genuinely lowers compute cost and latency in steady-state scenarios where capacity already matches demand.

  • ✗

    Use larger instance types to handle more requests per instance.

    Why it's wrong here

    Larger instances raise per-instance throughput yet still require the same multi-minute provisioning cycle to add capacity, leaving burst latency spikes unresolved. It tempts because vertical scaling is a valid cost-and-throughput lever when traffic is predictable and scaling events are infrequent.

  • ✓

    Configure the endpoint with a target tracking scaling policy and pre-warm additional instances during expected traffic surges.

    Why this is correct

    Target tracking maintains utilisation near a setpoint, but new instances still need minutes to provision and load the model. Pre-warming instances ahead of predicted surges removes that cold-start delay, so capacity exists before the burst arrives rather than scaling reactively.

  • ✗

    Set the endpoint to scale down slowly to maintain capacity.

    Why it's wrong here

    Slowing scale-down keeps idle capacity longer, which raises cost without accelerating scale-up, so burst spikes still hit cold capacity. It tempts because retaining warm instances does smooth demand fluctuations in scenarios where traffic dips are brief and cost is secondary.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.