MLA-C01 ML Model Development Practice Question
A team is fine-tuning a foundation model using reinforcement learning from human feedback (RLHF) on SageMaker. They have a dataset of human preferences. Which SageMaker capability is most suitable for the reward model training step?
⚠ Common exam trap
MLA-C01 often tests whether candidates confuse data-labeling (Ground Truth) or pre-built model hubs (JumpStart) with the custom training capability needed for RLHF reward models — the key is recognizing the need for a custom training script.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
SageMaker Training with a custom PyTorch container
Training a reward model for RLHF requires a custom training loop over human preference pairs, typically using a PyTorch model with a regression or ranking loss. SageMaker Training with a custom PyTorch container gives full control over the training script, loss function, and data pipeline needed for this step. The other SageMaker capabilities are higher-level or data-labeling services that do not provide this flexibility.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
SageMaker JumpStart
Why it's wrong here
JumpStart provides pre-trained models, notebooks and one-click deployment for common tasks; it does not train a reward model from preference pairs. It is tempting because it accelerates foundation-model work, and would be correct for deploying or fine-tuning a ready-made model, but the reward-model training step needs a custom training job.
- ✗
SageMaker Ground Truth
Why it's wrong here
Ground Truth labels raw data through human annotators and workers; it produces the preference dataset rather than training the reward model from it. It is tempting because RLHF depends on human preference data, and Ground Truth would be the right choice for collecting those comparisons, but the stem already supplies the dataset.
- ✗
SageMaker Autopilot
Why it's wrong here
Autopilot automates feature engineering and model selection for tabular supervised learning, not reward-model training on preference pairs. It is tempting because it builds models with minimal code, and would suit predicting a business metric from structured CSV data, but it cannot consume human preference comparisons or produce the scalar reward signal RLHF requires.
- ✓
SageMaker Training with a custom PyTorch container
Why this is correct
SageMaker Training with a custom PyTorch container suits reward-model training because RLHF reward models are bespoke regression networks, not standard supervised tasks. A custom container lets you define the pairwise preference loss, initialise from the fine-tuned foundation model's weights, and control hyperparameters — flexibility that built-in algorithms and Autopilot cannot provide for this scenario.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.