Courseiva
easyMultiple Choice

Configuring Auto-Scaling for SageMaker Real-Time Endpoints

A company has a trained machine learning model that needs to be deployed as a real-time inference endpoint on Amazon SageMaker. The endpoint must automatically scale based on incoming traffic. Which SageMaker feature should be used?

⚠ Common exam trap

Candidates often confuse SageMaker Elastic Inference (which accelerates inference) with auto scaling, or they assume Batch Transform can be used for real-time endpoints, but only Endpoint Auto Scaling directly manages dynamic instance count based on traffic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

SageMaker Endpoint Auto Scaling

Amazon SageMaker Endpoint Auto Scaling is the correct feature because it automatically adjusts the number of instances serving a real-time inference endpoint based on the incoming traffic load. It uses Application Auto Scaling policies, which monitor CloudWatch metrics (e.g., InvocationsPerInstance) to scale in or out, ensuring low latency and cost efficiency without manual intervention.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    SageMaker Endpoint Auto Scaling

    Why this is correct

    SageMaker Endpoint Auto Scaling dynamically adjusts the number of instances behind a real-time endpoint using target-tracking or step-scaling policies, driven by metrics such as InvocationsPerInstance. This directly satisfies the stem's requirement to scale automatically with incoming traffic, unlike batch transform or asynchronous inference, which cannot serve continuous real-time requests.

  • ✗

    SageMaker Elastic Inference

    Why it's wrong here

    Elastic Inference attaches fractional GPU acceleration to existing instances; it changes compute hardware, not instance count. It suits cost reduction for CPU-hosted inference, but scaling on incoming traffic needs an Application Auto Scaling policy targeting the endpoint variant, which Elastic Inference cannot provide.

  • ✗

    SageMaker Batch Transform

    Why it's wrong here

    Batch Transform processes entire datasets asynchronously in a single job, not as a persistent endpoint that can scale in real time to fluctuating traffic. It is tempting because it handles large-scale inference efficiently when latency is not critical, and would be correct for offline predictions on a static dataset where a real-time endpoint is unnecessary.

  • ✗

    SageMaker Model Monitor

    Why it's wrong here

    Model Monitor detects data drift and quality violations on deployed endpoints; it does not provision capacity or adjust instance counts. It would be the right choice for ongoing bias and drift detection, but automatic traffic-based scaling requires configuring an autoscaling policy on the endpoint's production variant.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on MLA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?

medium
  • A.Switch to batch transform for all inference requests.
  • B.Use a larger instance type to handle peak traffic.
  • ✓ C.Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.
  • D.Deploy multiple models behind an Application Load Balancer.

Why C: SageMaker endpoints support automatic scaling through target tracking scaling policies based on Amazon CloudWatch metrics like InvocationsPerInstance. This allows the endpoint to dynamically adjust the number of instances in response to real-time traffic spikes, scaling out when demand increases and scaling in when it decreases, which optimizes cost by only paying for the capacity needed at any given time.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.