Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A machine learning team is training a large…

A machine learning team is training a large natural language processing model on Amazon SageMaker using the SageMaker Hugging Face container. The training job runs on multiple instances and uses Managed Spot Training to reduce costs. However, the job frequently gets interrupted by Spot interruptions, causing long training times. What should the team do to mitigate this issue?

⚠ Common exam trap

MLA-C01 often tests the trade-off between cost savings and reliability with Spot instances — candidates pick 'disable Spot and use On-Demand' as the safe answer, but the question asks to mitigate interruptions while presumably keeping cost benefits, making checkpointing the correct choice.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable checkpointing and increase the number of save intervals

Managed Spot Training on SageMaker can be interrupted when EC2 reclaims Spot capacity. Enabling checkpointing (saving model state to Amazon S3 at intervals) allows training to resume from the last checkpoint instead of restarting from scratch. Increasing the number of save intervals reduces the amount of lost work per interruption, directly mitigating long training times.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a reserved capacity with Savings Plans

    Why it's wrong here

    Savings Plans and reserved capacity discount compute pricing but do not change the Spot capacity pool, so interruption frequency and checkpoint-based resumption remain identical. They suit steady, interruption-tolerant workloads where cost reduction is the goal, not scenarios requiring the job to survive Spot reclaim events.

  • ✗

    Use a larger instance type to finish faster

    Why it's wrong here

    A larger instance type does not reduce the frequency of Spot interruptions; it only shortens each checkpoint interval, so the job still restarts from the last checkpoint repeatedly. Larger instances are tempting because they speed up training, and would be correct if the bottleneck were compute capacity rather than interruption frequency.

  • ✓

    Enable checkpointing and increase the number of save intervals

    Why this is correct

    Checkpointing writes model state to Amazon S3 at each save interval, so when a Spot interruption reclaims an instance, the job resumes from the last checkpoint rather than restarting. Increasing save frequency reduces lost progress per interruption, directly mitigating the long training times caused by frequent Spot capacity reclamation.

  • ✗

    Disable Managed Spot Training and use On-Demand instances

    Why it's wrong here

    On-Demand instances eliminate the interruption mechanism entirely, but the stem's objective is mitigating interruptions while retaining Managed Spot Training's cost benefit. On-Demand is the right choice when training must run uninterrupted and cost is secondary, not when Spot savings are required.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.