Courseiva
ML Model Development →mediumMultiple Select

MLA-C01 ML Model Development Practice Question

A company is training a large NLP model on SageMaker and wants to reduce costs by using Spot Instances. Which TWO configurations should they implement to handle Spot interruptions gracefully?

⚠ Common exam trap

MLA-C01 often tests Spot Instance handling, and candidates may confuse max_wait with max_run or overlook the need for checkpointing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set `use_spot_instances=True` and `max_wait` in the estimator

Option B is correct because in the SageMaker SDK estimator you must explicitly set use_spot_instances=True to enable managed Spot training, and max_wait defines the maximum wall-clock time SageMaker will wait for the Spot capacity (including interruptions and restarts), which is required to let training resume after an interruption. Option E is correct because enabling checkpointing (e.g., via checkpoint_s3_uri) periodically saves model state to Amazon S3, so when a Spot instance is reclaimed the job can restart from the last checkpoint instead of from scratch, which is the core mechanism for handling interruptions gracefully. Option A is not correct because using a single large instance does not meaningfully reduce interruption probability and actually increases the cost/impact of a single interruption. Option C is not correct because increasing max_run only sets the maximum training duration; it does not help the job survive or recover from a Spot interruption. Option D is not correct because keep_alive_period is used for managed warm pools to reduce cold-start latency between jobs, not to preserve training state across Spot interruptions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a single large instance to reduce interruption probability

    Why it's wrong here

    A single large instance concentrates all training on one Spot capacity pool, so one reclaim terminates the whole job; spreading across multiple instances and pools raises the chance at least one survives. Single-instance setups suit steady On-Demand workloads, not interruption-tolerant Spot training.

  • ✓

    Set `use_spot_instances=True` and `max_wait` in the estimator

    Why this is correct

    Setting `use_spot_instances=True` with `max_wait` enables managed Spot training, where SageMaker checkpoints to Amazon S3 and resumes automatically after interruption, within the specified waiting window. This satisfies the stem's requirement to handle interruptions gracefully while cutting costs, since training continues rather than restarting from scratch.

  • ✗

    Increase the `max_run` parameter to allow longer training

    Why it's wrong here

    max_run caps total training duration; raising it lets a job run longer before timeout but provides no checkpointing or resumption when Spot capacity is reclaimed. It suits long On-Demand jobs needing a higher time limit, not graceful Spot interruption handling.

  • ✗

    Use `keep_alive_period` to keep the instance alive after training

    Why it's wrong here

    keep_alive_period applies to SageMaker inference endpoints or managed warm pools, retaining instances between jobs; it does nothing for a training job interrupted mid-run. It is the right setting when reducing cold-start latency for repeated inference or processing workloads.

  • ✓

    Enable checkpointing to save model state periodically

    Why this is correct

    Checkpointing periodically writes model artefacts to Amazon S3, so when a Spot Instance is reclaimed mid-training, a new instance resumes from the last saved state rather than restarting, satisfying the requirement to handle interruptions gracefully without losing completed training progress.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.