Courseiva
ML Model Development →mediumMultiple Select

MLA-C01 ML Model Development Practice Question

A machine learning engineer is using SageMaker Pipelines to automate the training and deployment of a model. The pipeline includes a processing step for feature engineering, a training step, and a model registration step. The engineer wants to ensure that the pipeline is reproducible and that the model artifacts are versioned. Which two actions should be taken? (Choose two.)

⚠ Common exam trap

Candidates often confuse tracking tools like SageMaker Experiments with actual versioning mechanisms, or assuming that caching provides reproducibility.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Store the training dataset in Amazon S3 with versioning enabled and reference the specific version in the pipeline.

Reproducibility requires pinning inputs, and versioning requires a registry for artifacts. Enabling S3 versioning and referencing a specific version ensures the training data is immutable and traceable. Registering the model in the SageMaker Model Registry creates a versioned entry that links to the artifacts and metadata. Together, these actions provide a reproducible pipeline and a versioned model catalog.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Configure the training step to use the latest SageMaker training image without specifying a version.

    Why it's wrong here

    Using the latest training image without a version tag introduces nondeterminism because the image could change over time, altering the training environment. For reproducibility, you should pin the image URI to a specific version or digest. This action would undermine reproducibility rather than support it, and it does not address model artifact versioning.

  • ✓

    Store the training dataset in Amazon S3 with versioning enabled and reference the specific version in the pipeline.

    Why this is correct

    Enabling S3 versioning and referencing a specific object version ensures that the exact dataset used for training is immutable and traceable. This supports reproducibility because rerunning the pipeline will use the same data version. Without versioning, the dataset could be overwritten, leading to inconsistent results. This action directly addresses the requirement for reproducibility and versioning of inputs.

  • ✓

    Register the trained model in the SageMaker Model Registry with a model package group and version.

    Why this is correct

    The SageMaker Model Registry provides a centralized catalog for model versions, each with metadata and approval status. Registering the model creates a versioned entry that links to the model artifacts in S3, enabling traceability and controlled deployment. This directly fulfills the requirement for model versioning and supports reproducibility by associating the model with its training data and parameters.

  • ✗

    Enable pipeline caching to reuse previous step outputs.

    Why it's wrong here

    Pipeline caching can speed up executions by reusing outputs when inputs and parameters are unchanged, but it does not provide versioning or reproducibility guarantees. It may even mask changes if not carefully managed. Caching is an optimization, not a versioning mechanism, and it does not ensure that model artifacts are cataloged with lineage.

  • ✗

    Use SageMaker Experiments to log the pipeline execution parameters and metrics.

    Why it's wrong here

    SageMaker Experiments helps track and compare runs, but it does not by itself guarantee reproducibility or versioning of model artifacts. It is a monitoring and analysis tool, not a mechanism for immutable data or model versioning. While useful for auditing, it does not satisfy the core need to pin data versions and register models with lineage.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.