Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A healthcare startup fine-tunes a model to generate patient education materials. They want to ensure the model never gives medical advice, only information. They add a safety instruction, but the model sometimes still gives advice. What advanced technique should they apply?
⚠ Common exam trap
Candidates often mistakenly believe that simple post-processing filters or static embedding comparisons are sufficient to enforce safety. However, only advanced alignment techniques like RLHF can truly align the model's generation, as it changes the model's behavior during training rather than applying brittle surface-level checks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply RLHF with a reward model that penalizes outputs containing medical advice
RLHF (Reinforcement Learning from Human Feedback) directly addresses the model's behavior by training a reward model that penalizes outputs containing medical advice. This aligns the model's generation with the safety instruction at a fundamental level, rather than relying on brittle post-hoc filters or static embeddings that can be easily circumvented by novel phrasings.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Hard-code a list of prohibited phrases in a post-processing script
Why it's wrong here
A hard-coded phrase list only catches exact known strings, missing paraphrased or novel advice, so it cannot guarantee the constraint. It suits blocking a small fixed set of banned terms; the scenario needs a semantic guardrail that generalises beyond literal matches.
- ✗
Add a secondary classifier to rewrite any detected advice into general information
Why it's wrong here
A classifier that rewrites detected advice into general information still permits advice to be generated first, and rewriting may preserve directive content. It suits style or tone transformation tasks; the scenario needs a guardrail that blocks or regenerates unsafe output before it reaches the user.
- ✗
Use semantic similarity to a 'medical advice' embedding and reject if close
Why it's wrong here
Embedding similarity thresholds flag text resembling medical advice but cannot reliably separate advice from information, and rejecting outputs degrades usefulness without correcting the model. It suits semantic search or deduplication; the scenario needs a guardrail that classifies and rewrites unsafe generations.
- ✓
Apply RLHF with a reward model that penalizes outputs containing medical advice
Why this is correct
RLHF trains a reward model that scores outputs, then optimises the model against it; penalising medical-advice content directly shapes generation away from advice. A single safety instruction only conditions the prompt, so it cannot reliably enforce the constraint across varied inputs.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.