Courseiva
ML Model Development →hardMultiple Select

MLA-C01 ML Model Development Practice Question

A data scientist is training a model with SageMaker and needs to reduce the cost of a long-running training job that can tolerate interruptions. The job uses a custom training script and reads data from Amazon S3. The data scientist wants the job to resume from the last saved state if the underlying compute is reclaimed. (Choose two.)

⚠ Common exam trap

The trap here is assuming spot training alone preserves progress, when checkpointing to Amazon S3 is what enables resuming after an interruption.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the estimator to use managed spot training and set max_wait time greater than max_run time.

Managed spot training lowers cost by using Spot Instances, and checkpointing to Amazon S3 lets a restarted job resume from the last saved state. Setting max_wait above max_run gives the job time to wait for capacity and to recover from interruptions. Network isolation, larger volumes, and restarting from scratch do not support the resume requirement and can increase cost or break access.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the estimator's enable_spot_training parameter to true and rely on SageMaker to automatically restart the job from the beginning after an interruption.

    Why it's wrong here

    The relevant parameter is train_use_spot_instances, not enable_spot_training, and automatic restart from the beginning would discard progress. Without checkpointing to Amazon S3, a restarted job cannot resume from the last state. This option misnames the parameter and describes behavior that does not reduce cost effectively.

  • ✗

    Increase the volume_size parameter so that the container's local disk can hold the entire dataset and all checkpoints during training.

    Why it's wrong here

    Volume size affects local disk capacity, not cost or interruptibility. Checkpoints must be written to Amazon S3 to survive an instance replacement; local disk is lost when the spot instance is reclaimed. Increasing volume size adds storage cost and does not enable resuming after an interruption.

  • ✓

    Configure the estimator to use managed spot training and set max_wait time greater than max_run time.

    Why this is correct

    Managed spot training uses Amazon EC2 Spot Instances, which are cheaper but can be interrupted. Setting max_wait greater than max_run allows the job to wait for capacity and to resume after an interruption within the overall wait window. This directly addresses cost reduction while tolerating interruptions, and it is a required configuration for the resume behavior to be useful.

  • ✗

    Enable network isolation on the estimator to prevent the training container from accessing the internet during spot interruptions.

    Why it's wrong here

    Network isolation restricts outbound network access, which would block the container from reaching Amazon S3 for data and checkpoints unless a VPC endpoint is configured. It does not help with spot interruptions or cost. Enabling it in this scenario would likely break data access and checkpointing rather than support resuming the job.

  • ✓

    Set the checkpoint_s3_uri parameter on the estimator to an Amazon S3 location where the training script saves checkpoints.

    Why this is correct

    Checkpointing writes model state to Amazon S3 during training. When a spot interruption occurs, SageMaker restarts the job and the script can load the latest checkpoint from that S3 location, resuming instead of starting over. Without this parameter, the job would restart from scratch, wasting the work already done and reducing the cost benefit.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.