MLA-C01 ML Model Development Practice Question
A company is training a large NLP model on SageMaker and wants to reduce costs by using Spot Instances. Which TWO configurations should they implement to handle Spot interruptions gracefully?
⚠ Common exam trap
MLA-C01 often tests Spot Instance handling, and candidates may confuse max_wait with max_run or overlook the need for checkpointing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set `use_spot_instances=True` and `max_wait` in the estimator
Option B is correct because in the SageMaker SDK estimator you must explicitly set use_spot_instances=True to enable managed Spot training, and max_wait defines the maximum wall-clock time SageMaker will wait for the Spot capacity (including interruptions and restarts), which is required to let training resume after an interruption. Option E is correct because enabling checkpointing (e.g., via checkpoint_s3_uri) periodically saves model state to Amazon S3, so when a Spot instance is reclaimed the job can restart from the last checkpoint instead of from scratch, which is the core mechanism for handling interruptions gracefully. Option A is not correct because using a single large instance does not meaningfully reduce interruption probability and actually increases the cost/impact of a single interruption. Option C is not correct because increasing max_run only sets the maximum training duration; it does not help the job survive or recover from a Spot interruption. Option D is not correct because keep_alive_period is used for managed warm pools to reduce cold-start latency between jobs, not to preserve training state across Spot interruptions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a single large instance to reduce interruption probability
Why it's wrong here
A single large instance concentrates all training on one Spot capacity pool, so one reclaim terminates the whole job; spreading across multiple instances and pools raises the chance at least one survives. Single-instance setups suit steady On-Demand workloads, not interruption-tolerant Spot training.
- ✓
Set `use_spot_instances=True` and `max_wait` in the estimator
Why this is correct
Setting `use_spot_instances=True` with `max_wait` enables managed Spot training, where SageMaker checkpoints to Amazon S3 and resumes automatically after interruption, within the specified waiting window. This satisfies the stem's requirement to handle interruptions gracefully while cutting costs, since training continues rather than restarting from scratch.
- ✗
Increase the `max_run` parameter to allow longer training
Why it's wrong here
max_run caps total training duration; raising it lets a job run longer before timeout but provides no checkpointing or resumption when Spot capacity is reclaimed. It suits long On-Demand jobs needing a higher time limit, not graceful Spot interruption handling.
- ✗
Use `keep_alive_period` to keep the instance alive after training
Why it's wrong here
keep_alive_period applies to SageMaker inference endpoints or managed warm pools, retaining instances between jobs; it does nothing for a training job interrupted mid-run. It is the right setting when reducing cold-start latency for repeated inference or processing workloads.
- ✓
Enable checkpointing to save model state periodically
Why this is correct
Checkpointing periodically writes model artefacts to Amazon S3, so when a Spot Instance is reclaimed mid-training, a new instance resumes from the last saved state rather than restarting, satisfying the requirement to handle interruptions gracefully without losing completed training progress.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.