MLA-C01 ML Model Development Practice Question
A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. The training job is checkpointed every 5 minutes. However, the job gets interrupted frequently and never completes. What is the MOST likely cause?
⚠ Common exam trap
MLA-C01 often tests whether candidates blame configuration errors (S3 URI, retries) instead of the fundamental trade-off between checkpoint frequency and interruption rate — the trap is missing that frequent interruptions require more frequent checkpoints.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The checkpoint interval is too long relative to the interruption frequency
SageMaker spot instances can be interrupted with little notice. If the checkpoint interval (5 minutes) is longer than the average time between interruptions, the job loses more progress than it saves, so it never completes. The most likely cause is that the checkpoint interval is too long relative to the interruption frequency, preventing effective resume.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The checkpoint interval is too long relative to the interruption frequency
Why this is correct
Spot capacity reclaims instances faster than the five-minute checkpoint cadence, so each interruption discards up to five minutes of progress before the next save. Shortening the checkpoint interval relative to interruption frequency satisfies the stem's requirement that the job actually complete.
- ✗
The checkpoint S3 URI is incorrect
Why it's wrong here
An incorrect checkpoint S3 URI prevents checkpoints being written or resumed, so each spot reclaim restarts from scratch and the job never finishes. It is tempting because a bad path seems like a configuration error, but the URI itself does not cause the interruptions; it only prevents recovery from them.
- ✗
The instance type is too small for the training job
Why it's wrong here
An undersized instance causes slow training or out-of-memory failures, not repeated spot interruptions; capacity reclaims come from spot availability. It is tempting because resource exhaustion also stops jobs, but the frequent-interruption symptom points to checkpointing or retry configuration, not instance sizing.
- ✗
The job is configured with too few max retries
Why it's wrong here
Too few max retries means the job stops after a small number of spot reclaims instead of resuming from the last checkpoint. It is tempting because retries relate to interruptions, but the setting governs how many times SageMaker restarts, and the frequent interruptions themselves stem from spot capacity.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.