Courseiva
hardMultiple ChoiceObjective-mapped

Configuring Auto-Scaling for SageMaker Real-Time Endpoints

A financial services company is deploying a real-time fraud detection model using Amazon SageMaker. The model is a gradient boosting model (XGBoost) trained on historical transaction data. The inference endpoint uses an ml.m5.2xlarge instance with a single variant. Recently, the company has experienced a 3x increase in transaction volume during peak hours, causing inference latency to exceed the 200ms SLA. The data science team has already optimized the model by reducing the number of trees and feature set, but the latency remains high during spikes. The team considers using SageMaker's built-in scaling policies. They currently have a single endpoint with one production variant. The team wants to maintain low latency without over-provisioning resources. They have ruled out model changes. Which approach should the team take?

Quick Answer

The correct choice targets the metric that actually reflects load on each instance behind an endpoint: invocations per instance, rather than something like CPU, which can behave inconsistently across model types and instance sizes. A target tracking policy built on SageMakerVariantInvocationsPerInstance lets Application Auto Scaling continuously compare the observed value to a target and add or remove instances automatically, so the fleet grows during the 3x peak-hour surge and shrinks back down once volume returns to normal, which is exactly what maintaining low latency without over-provisioning calls for. This is a better fit than simply adding a fixed number of extra instances or relying on manual scaling, because those approaches either waste capacity during quiet periods or fail to react fast enough when a spike arrives, and it's a better fit than further model tuning since the team already ruled that out and had already reduced trees and features without solving the spike-driven latency problem. The underlying pattern to recognize is that inference latency problems caused by variable traffic volume, rather than by the model itself, are solved by automatically adjusting instance count to match invocation load, and target tracking on an invocations-per-instance metric is SageMaker's purpose-built mechanism for doing that without manual thresholds or guesswork.

⚠ Common exam trap

A common mix-up: candidates confuse scheduled scaling (Option C) as a valid solution for predictable peaks, but the question's emphasis on 'real-time' and 'without over-provisioning' points to dynamic scaling, which target tracking provides; scheduled scaling cannot adapt to unexpected volume variations within the peak window.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Configure an Application Auto Scaling target tracking scaling policy for the variant based on the 'SageMakerVariantInvocationsPerInstance' metric, with a target value that keeps the inference latency within the SLA.

SageMaker's built-in target tracking scaling policy using the 'SageMakerVariantInvocationsPerInstance' metric allows the endpoint to automatically adjust the instance count based on real-time invocation load. By setting a target value that correlates with the 200ms SLA, the policy dynamically scales out during traffic spikes and scales in during lulls, preventing over-provisioning while maintaining low latency. This approach directly addresses the 3x peak-hour volume increase without requiring manual intervention or model changes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Configure an Application Auto Scaling target tracking scaling policy for the variant based on the 'SageMakerVariantInvocationsPerInstance' metric, with a target value that keeps the inference latency within the SLA.

    Why this is correct

    This auto-scales based on load.

  • Deploy the model on multiple endpoints behind an Application Load Balancer.

    Why it's wrong here

    Adds complexity and latency.

  • Use scheduled scaling to increase the instance count during known peak hours.

    Why it's wrong here

    Scheduled scaling cannot handle unexpected spikes.

  • Manually increase the instance count during peak hours.

    Why it's wrong here

    Manual scaling is not automated.

About these practice questions

One of 835 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

3 more ways this is tested on MLA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?

medium
  • A.Switch to batch transform for all inference requests.
  • B.Use a larger instance type to handle peak traffic.
  • C.Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.
  • D.Deploy multiple models behind an Application Load Balancer.

Why C: SageMaker endpoints support automatic scaling through target tracking scaling policies based on Amazon CloudWatch metrics like InvocationsPerInstance. This allows the endpoint to dynamically adjust the number of instances in response to real-time traffic spikes, scaling out when demand increases and scaling in when it decreases, which optimizes cost by only paying for the capacity needed at any given time.

Variation 2. A company has a trained machine learning model that needs to be deployed as a real-time inference endpoint on Amazon SageMaker. The endpoint must automatically scale based on incoming traffic. Which SageMaker feature should be used?

easy
  • A.SageMaker Endpoint Auto Scaling
  • B.SageMaker Elastic Inference
  • C.SageMaker Batch Transform
  • D.SageMaker Model Monitor

Why A: Amazon SageMaker Endpoint Auto Scaling is the correct feature because it automatically adjusts the number of instances serving a real-time inference endpoint based on the incoming traffic load. It uses Application Auto Scaling policies, which monitor CloudWatch metrics (e.g., InvocationsPerInstance) to scale in or out, ensuring low latency and cost efficiency without manual intervention.

Variation 3. A company is deploying a real-time inference endpoint for a natural language processing model using Amazon SageMaker. The model is a fine-tuned BERT variant. The endpoint has been running for two weeks with acceptable latency (average 200 ms). However, over the past 24 hours, the latency has increased to an average of 800 ms, and the number of simultaneous requests has doubled. The team expects traffic to continue to grow. The current endpoint configuration uses a single ml.m5.large instance. The model is loaded into memory once, and the inference framework is PyTorch. The team needs to maintain latency under 500 ms. Which course of action should the team take to address the latency increase while minimizing cost?

easy
  • A.Switch to ml.c5.large instances because CPU-optimized instances provide better inference performance for NLP models.
  • B.Increase the instance size to ml.m5.xlarge and keep a single instance.
  • C.Enable automatic scaling for the endpoint with a target average latency of 500 ms and use multiple ml.m5.large instances.
  • D.Implement a multi-model endpoint with multiple ml.m5.large instances and use Amazon Elastic Inference (EI) accelerators.

Why C: The latency increase is caused by a doubling of simultaneous requests overwhelming a single ml.m5.large instance. Enabling automatic scaling with a target average latency of 500 ms allows SageMaker to add more ml.m5.large instances as traffic grows, distributing the load and keeping latency under the threshold. This approach minimizes cost by scaling only when needed, rather than over-provisioning a larger instance.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.