Courseiva
hardMultiple Choice

MLA-C01 Target tracking scaling policy Practice Question

A financial services company uses Amazon SageMaker to deploy a fraud detection model for real-time inference. The model is deployed on an ml.m5.large instance with a SageMaker real-time endpoint. The endpoint has an auto scaling policy configured using a custom scaling policy based on average CPU utilization, with scale out threshold at 70% and scale in threshold at 30%. During a flash sale event, the traffic to the endpoint spikes tenfold within minutes. The endpoint fails to handle the load, resulting in increased latency and timeouts. The data science team needs to improve the scalability of the endpoint to handle sudden traffic spikes. Which solution should the team implement?

⚠ Common exam trap

MLA-C01 often tests the misconception that CPU-based scaling is always sufficient, when for inference endpoints invocation-based target tracking is more responsive to traffic spikes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Replace the custom scaling policy with a target tracking scaling policy based on the number of invocations per instance, with a target value of 1000.

A target tracking scaling policy based on invocations per instance directly ties scaling to the actual workload metric for a real-time endpoint, allowing faster and more accurate scale-out during sudden traffic spikes. Unlike CPU-based custom scaling, invocation-based target tracking reacts to request volume, which is the true driver of load for inference. This improves the endpoint's ability to handle a tenfold spike.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Implement a SageMaker Model Ensemble with two additional models to balance the load.

    Why it's wrong here

    An ensemble runs two additional models per request, multiplying compute cost and latency while still executing on the same fixed endpoint capacity, so it cannot absorb tenfold traffic. Ensembles suit improving prediction accuracy through model diversity, not horizontal scaling under sudden load.

  • ✓

    Replace the custom scaling policy with a target tracking scaling policy based on the number of invocations per instance, with a target value of 1000.

    Why this is correct

    Invocation-based target tracking scales on request volume, the metric that actually surges tenfold during flash sales, whereas CPU lags behind sudden concurrency spikes. It satisfies the sudden-traffic requirement by adding instances before latency degrades, rather than reacting to already-saturated compute.

  • ✗

    Implement a SageMaker Inference Pipeline with a pre-processing step to reduce model input size.

    Why it's wrong here

    A pre-processing step inside an Inference Pipeline reduces payload size and feature engineering latency; it does not add endpoint instances, so tenfold concurrent request volume still queues against the same ml.m5.large capacity. Pipelines suit chaining preprocessing with inference, not absorbing sudden traffic spikes.

  • ✗

    Switch to a GPU instance type, such as ml.p3.2xlarge, to increase compute capacity.

    Why it's wrong here

    A GPU instance accelerates per-request matrix computation but leaves a single instance behind the endpoint, so tenfold concurrent traffic still saturates it; the CPU-based target-tracking policy also scales on utilisation, not request backlog. GPU instances suit compute-heavy deep learning inference, not burst concurrency.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.