Courseiva
ML Model Development →hardMultiple Select

MLA-C01 ML Model Development Practice Question

A company wants to use SageMaker to fine-tune a foundation model for a text generation task using RLHF (Reinforcement Learning from Human Feedback). Which THREE components are required in the RLHF pipeline?

⚠ Common exam trap

MLA-C01 often tests whether candidates confuse RLHF components with general fine-tuning techniques like LoRA or with GAN-style discriminators, causing them to pick optional or unrelated options.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A pre-trained base model

Option B is correct because RLHF always starts from a pre-trained foundation model (the policy) that already has broad language capabilities; this base model is then fine-tuned with human preference signals rather than trained from scratch. Option D is correct because RLHF requires a reward model trained on human preference comparisons (e.g., chosen vs. rejected responses) to serve as a differentiable proxy for human judgment during optimization. Option E is correct because the policy is optimized against that reward model using a reinforcement learning algorithm such as PPO (Proximal Policy Optimization), typically with a KL penalty to the reference model to prevent reward hacking. Option A is not required: LoRA is a parameter-efficient fine-tuning technique that can optionally be used, but RLHF does not mandate adapters. Option C is not required: distinguishing generated from real text is a GAN discriminator concept, not part of the RLHF pipeline, which relies on preference-based reward modeling instead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    A LoRA adapter for parameter-efficient fine-tuning

    Why it's wrong here

    LoRA adapters reduce trainable parameters during supervised fine-tuning, but RLHF's required pipeline components are a reward model, a policy, and a PPO optimiser. LoRA is tempting as an efficiency technique, and would be correct when compute-constrained supervised fine-tuning is the goal rather than preference-based optimisation.

  • ✓

    A pre-trained base model

    Why this is correct

    RLHF starts from a pre-trained base model, which supplies the generative prior that the policy is later refined from. Without it, there is no initial language model to sample responses from or to update during reinforcement learning, so the pipeline cannot begin.

  • ✗

    A classifier to distinguish generated text from real text

    Why it's wrong here

    RLHF uses a reward model trained on human preference rankings, not a real-versus-generated classifier. A classifier is tempting because GANs and adversarial training employ discriminators, but RLHF's reward model scores candidate responses against human preferences rather than detecting synthetic text.

  • ✓

    A reward model trained on human preferences

    Why this is correct

    The reward model encodes human preferences as a scalar signal, learned from ranked comparisons. It replaces a hand-crafted reward function, letting PPO optimise the policy toward outputs humans actually prefer, which is the defining mechanism of RLHF.

  • ✓

    A reinforcement learning algorithm such as PPO

    Why this is correct

    PPO supplies the policy-gradient update that optimises the language model against the reward model's scalar signal, satisfying RLHF's reinforcement learning stage. Without it, the pipeline reduces to supervised fine-tuning plus preference modelling, leaving the reward model's scores unused for policy improvement. PPO's clipped objective also stabilises training on this high-dimensional text generation task.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on MLA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company is fine-tuning a large language model using reinforcement learning from human feedback (RLHF). Which THREE components are typically required?

hard
  • A.A discriminative classifier
  • ✓ B.A reference model
  • ✓ C.A reward model
  • ✓ D.A policy model (the LLM)
  • E.A value function

Why B: In the standard RLHF pipeline, the policy model (D) is the large language model being fine-tuned; it generates responses and is updated via reinforcement learning (typically PPO) to maximize reward. A reward model (C) is trained on human preference comparisons to output a scalar score predicting human preference, and it supplies the reward signal that guides the policy's optimization. A reference model (B) is a frozen copy of the pre-RLHF model used to compute a KL-divergence penalty, keeping the policy from drifting too far from the original model and collapsing into degenerate, high-reward outputs. The other options are not required components: a discriminative classifier (A) is not part of RLHF (the reward model is a regression-style preference predictor, not a classifier), and a value function (E) is only an internal component of the PPO algorithm used to estimate advantages, not a standalone required model in the RLHF architecture.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.