A healthcare startup fine-tunes a model to generate patient education materials. They want to ensure the model never gives medical advice, only information. They add a safety instruction, but the model sometimes still gives advice. What advanced technique should they apply?
RLHF trains a reward model that scores outputs, then optimises the model against it; penalising medical-advice content directly shapes generation away from advice. A single safety instruction only conditions the prompt, so it cannot reliably enforce the constraint across varied inputs.
Why this answer
RLHF (Reinforcement Learning from Human Feedback) directly addresses the model's behavior by training a reward model that penalizes outputs containing medical advice. This aligns the model's generation with the safety instruction at a fundamental level, rather than relying on brittle post-hoc filters or static embeddings that can be easily circumvented by novel phrasings.
Exam trap
Candidates often mistakenly believe that simple post-processing filters or static embedding comparisons are sufficient to enforce safety. However, only advanced alignment techniques like RLHF can truly align the model's generation, as it changes the model's behavior during training rather than applying brittle surface-level checks.
How to eliminate wrong answers
Option A is wrong because hard-coding a list of prohibited phrases is brittle and fails against adversarial or paraphrased advice that doesn't match the exact phrases. Option B is wrong because adding a secondary classifier to rewrite detected advice introduces latency, potential for semantic drift, and cannot handle nuanced contexts where advice is implied rather than explicit. Option C is wrong because semantic similarity to a static 'medical advice' embedding is threshold-dependent and can produce false positives (flagging general information) or false negatives (missing advice phrased differently), and it does not train the model to avoid the behavior.