Courseiva

MLA-C01 Deployment and Orchestration of ML Workflows Practice Question

A media company runs a real-time recommendation model on a SageMaker endpoint. Traffic triples every evening between 18:00 and 22:00 and drops to near zero overnight, and the team wants to cut costs without underserving evening users. They want the endpoint to scale out automatically as invocations rise and scale back in when traffic falls. Which solution should they implement?

⚠ Common exam trap

The trap here is assuming Serverless Inference is always the cheapest elastic option, when a fixed maximum concurrency reintroduces the same idle cost as a permanently over-provisioned endpoint.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable automatic scaling on the endpoint with a target-tracking scaling policy based on the SageMakerVariantInvocationsPerInstance metric.

Target-tracking automatic scaling based on SageMakerVariantInvocationsPerInstance is the native mechanism for elastic real-time endpoints. It registers the variant as a scalable target with Application Auto Scaling, then adds or removes instances to hold the metric near the target. This directly matches a recurring evening peak followed by an overnight trough while keeping latency low, unlike fixed serverless concurrency or static over-provisioning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create a second endpoint and use an Application Load Balancer to route evening traffic between the two endpoints.

    Why it's wrong here

    An ALB cannot perform capacity-based scaling of SageMaker endpoints, and a second endpoint with static instance counts does not respond to invocation volume. You would still be paying for both endpoints all day, and routing alone does not add instances as load rises. This adds cost and operational complexity without solving the dynamic scaling requirement.

  • ✗

    Deploy the model to a SageMaker Serverless Inference endpoint and set the maximum concurrency to a fixed value equal to the evening peak.

    Why it's wrong here

    Serverless Inference scales automatically, but a fixed maximum concurrency set to the evening peak removes the cost benefit, since overnight traffic still reserves the same capacity ceiling and you pay for provisioned concurrency. It also adds cold-start latency after idle periods, which is undesirable for evening interactive recommendations. Serverless is better suited to intermittent, unpredictable traffic, not a predictable daily peak.

  • ✓

    Enable automatic scaling on the endpoint with a target-tracking scaling policy based on the SageMakerVariantInvocationsPerInstance metric.

    Why this is correct

    Target tracking on SageMakerVariantInvocationsPerInstance lets Application Auto Scaling add instances when per-instance invocation load exceeds the target and remove them when it drops, matching the nightly peak-and-trough pattern without manual intervention. This is the documented approach for provisioning real-time endpoint capacity dynamically, and it preserves latency by scaling out before instances saturate.

  • ✗

    Increase the instance count of the endpoint to the evening peak permanently and rely on Savings Plans to offset the extra cost.

    Why it's wrong here

    Provisioning for the peak around the clock means paying for idle instances overnight, and Savings Plans only discount committed spend rather than eliminate it. It ignores the stated goal of reducing cost during low-traffic hours. Static over-provisioning also wastes quota and slows deployments, whereas the requirement is elastic capacity tied to invocation demand.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.