An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)
RLHF requires a reward model that scores model outputs against learned human preferences, supplying the scalar signal PPO maximises. Without it, no preference-based objective exists, so the fine-tuning loop cannot optimise the foundation model toward preferred responses.
Why this answer
Option A is correct because RLHF requires a reward model that has been trained on human preference data to serve as the scalar reward signal guiding policy optimization. Option C is correct because PPO (Proximal Policy Optimization) is the standard reinforcement learning algorithm used to update the policy (the language model) against the reward model while constraining updates via a KL penalty to the reference model. Option D is correct because a preference dataset containing human rankings (e.g., chosen vs. rejected responses) is the foundational input used to train the reward model in the first place.
Option B is not essential to the RLHF training workflow itself; a validation set is useful for evaluation but is not a required RLHF component. Option E is not essential because PEFT methods like LoRA are an optional efficiency technique for fine-tuning, not a required element of the RLHF pipeline.
Exam trap
MLA-C01 often tests the distinction between RLHF's essential components (preference data, reward model, PPO) and optional optimizations like PEFT or evaluation datasets, causing candidates to select non-essential items.