Courseiva
ML Model Development →hardMultiple Choice

MLA-C01 ML Model Development Practice Question

A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. The training job is checkpointed every 5 minutes. However, the job gets interrupted frequently and never completes. What is the MOST likely cause?

⚠ Common exam trap

MLA-C01 often tests whether candidates blame configuration errors (S3 URI, retries) instead of the fundamental trade-off between checkpoint frequency and interruption rate — the trap is missing that frequent interruptions require more frequent checkpoints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The checkpoint interval is too long relative to the interruption frequency

SageMaker spot instances can be interrupted with little notice. If the checkpoint interval (5 minutes) is longer than the average time between interruptions, the job loses more progress than it saves, so it never completes. The most likely cause is that the checkpoint interval is too long relative to the interruption frequency, preventing effective resume.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The checkpoint interval is too long relative to the interruption frequency

    Why this is correct

    Spot capacity reclaims instances faster than the five-minute checkpoint cadence, so each interruption discards up to five minutes of progress before the next save. Shortening the checkpoint interval relative to interruption frequency satisfies the stem's requirement that the job actually complete.

  • ✗

    The checkpoint S3 URI is incorrect

    Why it's wrong here

    An incorrect checkpoint S3 URI prevents checkpoints being written or resumed, so each spot reclaim restarts from scratch and the job never finishes. It is tempting because a bad path seems like a configuration error, but the URI itself does not cause the interruptions; it only prevents recovery from them.

  • ✗

    The instance type is too small for the training job

    Why it's wrong here

    An undersized instance causes slow training or out-of-memory failures, not repeated spot interruptions; capacity reclaims come from spot availability. It is tempting because resource exhaustion also stops jobs, but the frequent-interruption symptom points to checkpointing or retry configuration, not instance sizing.

  • ✗

    The job is configured with too few max retries

    Why it's wrong here

    Too few max retries means the job stops after a small number of spot reclaims instead of resuming from the last checkpoint. It is tempting because retries relate to interruptions, but the setting governs how many times SageMaker restarts, and the frequent interruptions themselves stem from spot capacity.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.