MLA-C01 Practice Question: ML Solution Monitoring, Maintenance, and Security
A company wants to reduce costs for a SageMaker real-time endpoint that receives predictable traffic patterns: high during business hours and low at night. The model is a small PyTorch model. Which cost-saving strategy is most suitable?
⚠ Common exam trap
The trap is over-engineering with multi-model endpoints or batch transform when the question explicitly states predictable traffic — scheduled auto-scaling is the textbook answer for predictable patterns.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure auto-scaling with a scheduled scaling policy to add instances during business hours and reduce at night
Scheduled auto-scaling is designed for predictable traffic patterns: you define a schedule to scale out during business hours and scale in at night, matching capacity to demand and minimizing cost. For a small PyTorch model on a real-time endpoint, this is the most cost-effective and operationally simple strategy.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a single large instance to handle peak load
Why it's wrong here
A single large instance sized for peak load runs continuously at peak capacity, so overnight idle periods still incur full charges. It is tempting because one instance simplifies management, but it defeats the cost-saving goal; the correct approach scales instances down when traffic drops.
- ✗
Use a multi-model endpoint with multiple models
Why it's wrong here
Multi-model endpoints reduce hosting costs by loading several models behind one container, not by scaling capacity to match predictable diurnal traffic. It is tempting because it lowers per-model hosting overhead, but a single small PyTorch model gains nothing from co-location; the requirement is scheduling-based scaling.
- ✓
Configure auto-scaling with a scheduled scaling policy to add instances during business hours and reduce at night
Why this is correct
Scheduled scaling aligns endpoint instance counts with the predictable daytime peak and nightly trough, so capacity is added only when needed. This directly addresses the stem's cost-reduction goal for a small PyTorch model on a real-time endpoint.
- ✗
Switch to batch transform jobs and run nightly
Why it's wrong here
Batch transform processes discrete datasets on demand, so it cannot serve the continuous, low-latency inference a real-time endpoint provides during business hours. It is tempting because nightly batch runs suit offline scoring, but this workload needs synchronous responses, which batch transform does not deliver.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.