Courseiva
ML Model Development →hardMultiple Select

MLA-C01 ML Model Development Practice Question

An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)

⚠ Common exam trap

MLA-C01 often tests the distinction between RLHF's essential components (preference data, reward model, PPO) and optional optimizations like PEFT or evaluation datasets, causing candidates to select non-essential items.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A reward model trained on the preference data

Option A is correct because RLHF requires a reward model that has been trained on human preference data to serve as the scalar reward signal guiding policy optimization. Option C is correct because PPO (Proximal Policy Optimization) is the standard reinforcement learning algorithm used to update the policy (the language model) against the reward model while constraining updates via a KL penalty to the reference model. Option D is correct because a preference dataset containing human rankings (e.g., chosen vs. rejected responses) is the foundational input used to train the reward model in the first place. Option B is not essential to the RLHF training workflow itself; a validation set is useful for evaluation but is not a required RLHF component. Option E is not essential because PEFT methods like LoRA are an optional efficiency technique for fine-tuning, not a required element of the RLHF pipeline.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    A reward model trained on the preference data

    Why this is correct

    RLHF requires a reward model that scores model outputs against learned human preferences, supplying the scalar signal PPO maximises. Without it, no preference-based objective exists, so the fine-tuning loop cannot optimise the foundation model toward preferred responses.

  • ✗

    A large validation dataset for final evaluation

    Why it's wrong here

    RLHF's essential components are a reward model, a policy optimisation algorithm and preference data; a large validation dataset supports final evaluation rather than the training loop itself. It tempts because held-out validation data is mandatory in conventional supervised fine-tuning, so engineers assume it carries over.

  • ✓

    The PPO (Proximal Policy Optimization) algorithm for model updates

    Why this is correct

    PPO performs the actual policy updates, adjusting the foundation model's weights to increase reward while constraining divergence from the previous policy. This clipped objective keeps RLHF training stable, which is why PPO is the standard optimiser in this workflow.

  • ✓

    A preference dataset with human rankings

    Why this is correct

    Human rankings supply the supervision RLHF depends on: they train the reward model and define what 'preferred' means. Without this preference dataset, no reward signal can be learned, so the policy has nothing to optimise against.

  • ✗

    A PEFT technique like LoRA

    Why it's wrong here

    LoRA is a parameter-efficient fine-tuning method that freezes base weights and trains low-rank adapters, which is not required for RLHF; RLHF needs a reward model, a policy optimiser such as PPO, and preference data. It tempts because PEFT is standard for supervised fine-tuning of large models on limited GPU memory.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.