MLA-C01 ML Model Development Practice Question
A company is fine-tuning a large language model using reinforcement learning from human feedback (RLHF). Which THREE components are typically required?
⚠ Common exam trap
The trap is including the value function (critic) as a required component because it appears in PPO; the question asks for the three architectural components of RLHF—policy, reward, and reference models.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A reference model
In the standard RLHF pipeline, the policy model (D) is the large language model being fine-tuned; it generates responses and is updated via reinforcement learning (typically PPO) to maximize reward. A reward model (C) is trained on human preference comparisons to output a scalar score predicting human preference, and it supplies the reward signal that guides the policy's optimization. A reference model (B) is a frozen copy of the pre-RLHF model used to compute a KL-divergence penalty, keeping the policy from drifting too far from the original model and collapsing into degenerate, high-reward outputs. The other options are not required components: a discriminative classifier (A) is not part of RLHF (the reward model is a regression-style preference predictor, not a classifier), and a value function (E) is only an internal component of the PPO algorithm used to estimate advantages, not a standalone required model in the RLHF architecture.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A discriminative classifier
Why it's wrong here
A discriminative classifier outputs a class label or probability, not the scalar preference score RLHF's reward model must produce from prompt-response pairs. It is tempting because classification and reward modelling share a backbone, but the reward model is trained on pairwise preference rankings, making a plain classifier the wrong head.
- ✓
A reference model
Why this is correct
RLHF needs a frozen reference model to compute the KL-divergence penalty against the fine-tuned policy, preventing reward hacking and keeping outputs close to the original pretrained distribution. This satisfies the requirement for a stable baseline during PPO optimisation.
- ✓
A reward model
Why this is correct
The reward model, trained on human preference comparisons, supplies the scalar reward signal that PPO uses to update the policy. Without it, RLHF has no learned objective to optimise, so it is a required component.
- ✓
A policy model (the LLM)
Why this is correct
The policy model is the LLM being optimised, generating responses that the reward model scores during RLHF. It satisfies the stem's requirement for a trainable component whose weights are updated via policy-gradient methods against reward signals, forming the core actor in the actor-critic setup alongside the reward and value models.
- ✗
A value function
Why it's wrong here
The value function estimates expected return for a policy state and belongs to the PPO optimiser, not the RLHF pipeline's required components. It is tempting because PPO is the usual RLHF algorithm, but the three components are the pre-trained model, the reward model and the reinforcement learning policy.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.