MLA-C01 ML Model Development Practice Question
An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)
⚠ Common exam trap
MLA-C01 often tests the distinction between RLHF's essential components (preference data, reward model, PPO) and optional optimizations like PEFT or evaluation datasets, causing candidates to select non-essential items.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A reward model trained on the preference data
Option A is correct because RLHF requires a reward model that has been trained on human preference data to serve as the scalar reward signal guiding policy optimization. Option C is correct because PPO (Proximal Policy Optimization) is the standard reinforcement learning algorithm used to update the policy (the language model) against the reward model while constraining updates via a KL penalty to the reference model. Option D is correct because a preference dataset containing human rankings (e.g., chosen vs. rejected responses) is the foundational input used to train the reward model in the first place. Option B is not essential to the RLHF training workflow itself; a validation set is useful for evaluation but is not a required RLHF component. Option E is not essential because PEFT methods like LoRA are an optional efficiency technique for fine-tuning, not a required element of the RLHF pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
A reward model trained on the preference data
Why this is correct
RLHF requires a reward model that scores model outputs against learned human preferences, supplying the scalar signal PPO maximises. Without it, no preference-based objective exists, so the fine-tuning loop cannot optimise the foundation model toward preferred responses.
- ✗
A large validation dataset for final evaluation
Why it's wrong here
RLHF's essential components are a reward model, a policy optimisation algorithm and preference data; a large validation dataset supports final evaluation rather than the training loop itself. It tempts because held-out validation data is mandatory in conventional supervised fine-tuning, so engineers assume it carries over.
- ✓
The PPO (Proximal Policy Optimization) algorithm for model updates
Why this is correct
PPO performs the actual policy updates, adjusting the foundation model's weights to increase reward while constraining divergence from the previous policy. This clipped objective keeps RLHF training stable, which is why PPO is the standard optimiser in this workflow.
- ✓
A preference dataset with human rankings
Why this is correct
Human rankings supply the supervision RLHF depends on: they train the reward model and define what 'preferred' means. Without this preference dataset, no reward signal can be learned, so the policy has nothing to optimise against.
- ✗
A PEFT technique like LoRA
Why it's wrong here
LoRA is a parameter-efficient fine-tuning method that freezes base weights and trains low-rank adapters, which is not required for RLHF; RLHF needs a reward model, a policy optimiser such as PPO, and preference data. It tempts because PEFT is standard for supervised fine-tuning of large models on limited GPU memory.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.