MLA-C01 ML Model Development Practice Question
A machine learning engineer is using SageMaker Pipelines to automate the training and deployment of a model. The pipeline includes a processing step for feature engineering, a training step, and a model registration step. The engineer wants to ensure that the pipeline is reproducible and that the model artifacts are versioned. Which two actions should be taken? (Choose two.)
⚠ Common exam trap
Candidates often confuse tracking tools like SageMaker Experiments with actual versioning mechanisms, or assuming that caching provides reproducibility.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store the training dataset in Amazon S3 with versioning enabled and reference the specific version in the pipeline.
Reproducibility requires pinning inputs, and versioning requires a registry for artifacts. Enabling S3 versioning and referencing a specific version ensures the training data is immutable and traceable. Registering the model in the SageMaker Model Registry creates a versioned entry that links to the artifacts and metadata. Together, these actions provide a reproducible pipeline and a versioned model catalog.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure the training step to use the latest SageMaker training image without specifying a version.
Why it's wrong here
Using the latest training image without a version tag introduces nondeterminism because the image could change over time, altering the training environment. For reproducibility, you should pin the image URI to a specific version or digest. This action would undermine reproducibility rather than support it, and it does not address model artifact versioning.
- ✓
Store the training dataset in Amazon S3 with versioning enabled and reference the specific version in the pipeline.
Why this is correct
Enabling S3 versioning and referencing a specific object version ensures that the exact dataset used for training is immutable and traceable. This supports reproducibility because rerunning the pipeline will use the same data version. Without versioning, the dataset could be overwritten, leading to inconsistent results. This action directly addresses the requirement for reproducibility and versioning of inputs.
- ✓
Register the trained model in the SageMaker Model Registry with a model package group and version.
Why this is correct
The SageMaker Model Registry provides a centralized catalog for model versions, each with metadata and approval status. Registering the model creates a versioned entry that links to the model artifacts in S3, enabling traceability and controlled deployment. This directly fulfills the requirement for model versioning and supports reproducibility by associating the model with its training data and parameters.
- ✗
Enable pipeline caching to reuse previous step outputs.
Why it's wrong here
Pipeline caching can speed up executions by reusing outputs when inputs and parameters are unchanged, but it does not provide versioning or reproducibility guarantees. It may even mask changes if not carefully managed. Caching is an optimization, not a versioning mechanism, and it does not ensure that model artifacts are cataloged with lineage.
- ✗
Use SageMaker Experiments to log the pipeline execution parameters and metrics.
Why it's wrong here
SageMaker Experiments helps track and compare runs, but it does not by itself guarantee reproducibility or versioning of model artifacts. It is a monitoring and analysis tool, not a mechanism for immutable data or model versioning. While useful for auditing, it does not satisfy the core need to pin data versions and register models with lineage.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.