Courseiva

MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security

A company deploys a real-time inference endpoint with auto-scaling using a target tracking policy based on average Invocations per instance. They notice that during a traffic spike, the endpoint scales out too late, causing increased latency. They want to scale proactively before the spike. Which strategy should they implement?

⚠ Common exam trap

MLA-C01 often tests reactive versus proactive scaling, so candidates pick target tracking variants or provisioned concurrency, missing that scheduled scaling is the only truly proactive option for predictable spikes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a scheduled scaling action to add capacity before the expected spike

A scheduled scaling action adds capacity at a predetermined time before the expected traffic spike, allowing the endpoint to be ready proactively. Target tracking reacts after metrics breach thresholds, which is inherently reactive and too slow for sharp spikes. Scheduled scaling is the correct strategy when spikes are predictable (e.g., known business hours or events).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable provisioned concurrency on the endpoint

    Why it's wrong here

    Provisioned concurrency pre-initialises execution environments to cut cold-start latency, but it fixes capacity rather than expanding it, so the endpoint still cannot add instances before the spike. It suits latency-sensitive steady workloads, not proactive scale-out driven by forecast demand.

  • ✗

    Pre-warm the endpoint by sending dummy requests

    Why it's wrong here

    Dummy requests consume capacity without predicting demand, so scaling still reacts after real traffic arrives. Pre-warming suits steady baseline throughput or cold-start mitigation, not anticipating a spike; scheduled scaling or a CloudWatch-based predictive policy triggers scale-out ahead of the surge.

  • ✓

    Use a scheduled scaling action to add capacity before the expected spike

    Why this is correct

    Scheduled scaling adds capacity at predetermined times, so instances are already running before the expected traffic spike. This proactive approach avoids the lag inherent in target tracking, which reacts only after Invocations per instance rises.

  • ✗

    Switch to a step scaling policy with a higher cooldown period

    Why it's wrong here

    A longer cooldown delays subsequent scaling actions, worsening late scale-out during a spike. Step scaling adjusts capacity on CloudWatch alarm thresholds and suits known, discrete load tiers, but it still reacts to breached thresholds rather than scaling ahead of forecast demand.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.