Courseiva
ML Model Development →hardMultiple Choice

MLA-C01 ML Model Development Practice Question

A team is fine-tuning a foundation model using reinforcement learning from human feedback (RLHF) on SageMaker. They have a dataset of human preferences. Which SageMaker capability is most suitable for the reward model training step?

⚠ Common exam trap

MLA-C01 often tests whether candidates confuse data-labeling (Ground Truth) or pre-built model hubs (JumpStart) with the custom training capability needed for RLHF reward models — the key is recognizing the need for a custom training script.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

SageMaker Training with a custom PyTorch container

Training a reward model for RLHF requires a custom training loop over human preference pairs, typically using a PyTorch model with a regression or ranking loss. SageMaker Training with a custom PyTorch container gives full control over the training script, loss function, and data pipeline needed for this step. The other SageMaker capabilities are higher-level or data-labeling services that do not provide this flexibility.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    SageMaker JumpStart

    Why it's wrong here

    JumpStart provides pre-trained models, notebooks and one-click deployment for common tasks; it does not train a reward model from preference pairs. It is tempting because it accelerates foundation-model work, and would be correct for deploying or fine-tuning a ready-made model, but the reward-model training step needs a custom training job.

  • ✗

    SageMaker Ground Truth

    Why it's wrong here

    Ground Truth labels raw data through human annotators and workers; it produces the preference dataset rather than training the reward model from it. It is tempting because RLHF depends on human preference data, and Ground Truth would be the right choice for collecting those comparisons, but the stem already supplies the dataset.

  • ✗

    SageMaker Autopilot

    Why it's wrong here

    Autopilot automates feature engineering and model selection for tabular supervised learning, not reward-model training on preference pairs. It is tempting because it builds models with minimal code, and would suit predicting a business metric from structured CSV data, but it cannot consume human preference comparisons or produce the scalar reward signal RLHF requires.

  • ✓

    SageMaker Training with a custom PyTorch container

    Why this is correct

    SageMaker Training with a custom PyTorch container suits reward-model training because RLHF reward models are bespoke regression networks, not standard supervised tasks. A custom container lets you define the pairwise preference loss, initialise from the fine-tuned foundation model's weights, and control hyperparameters — flexibility that built-in algorithms and Autopilot cannot provide for this scenario.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.