Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A company is deploying a generative AI system that generates customer-facing emails. The system must ensure outputs are not toxic, biased, or harmful. Which TWO techniques are most effective for reducing toxicity in model outputs without significantly affecting performance?
⚠ Common exam trap
A common misconception is that simply adjusting generation parameters (like temperature or token count) or adding more examples can effectively control toxicity, when in fact these methods do not address the root cause of harmful patterns in the model's behavior.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Fine-tune the model on a dataset of safe emails using reinforcement learning from human feedback (RLHF).
Option C is correct because fine-tuning with RLHF directly aligns the model's behavior with human preferences for safe, non-toxic content, teaching it to avoid biased or harmful outputs at the model level without degrading overall generation quality. Option D is correct because a toxicity detection and filtering layer such as Vertex AI Safety Filters evaluates generated text against configurable harm categories and blocks or suppresses toxic outputs before they reach customers, adding a reliable guardrail. Option A is not effective because increasing max output tokens only changes length limits and can actually introduce more opportunities for harmful content. Option B is not effective because a very low temperature reduces randomness but does not remove learned toxic patterns and can hurt creativity and diversity. Option E is not effective because 50 few-shot examples consume significant context, increase cost and latency, and provide only shallow, prompt-level guidance that is easily bypassed.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the maximum output token count to allow more context.
Why it's wrong here
Raising the token ceiling only lengthens responses; it applies no filtering or alignment signal, so toxic content can still be produced. It is tempting because larger output limits genuinely help tasks needing long-form context, such as summarising lengthy documents, but toxicity reduction requires moderation or safety tuning instead.
- ✗
Set temperature to a very low value (e.g., 0.1).
Why it's wrong here
Low temperature reduces sampling randomness, making outputs repetitive and deterministic, yet a biased or toxic distribution is still selected from. It is tempting because low temperature genuinely suits tasks demanding consistent, factual responses such as data extraction, but it does not filter harmful content; moderation or safety alignment does.
- ✓
Fine-tune the model on a dataset of safe emails using reinforcement learning from human feedback (RLHF).
Why this is correct
RLHF fine-tuning adjusts the model's weights toward human-preferred, non-toxic responses, reducing harmful generations at the source. Because the change is learned rather than bolted on, it preserves general capability better than prompt-only or post-hoc filtering approaches.
- ✓
Apply a toxicity detection and filtering layer using Vertex AI Safety Filters.
Why this is correct
Vertex AI Safety Filters intercept both prompts and responses, scoring them against configurable harm categories such as toxicity, hate speech and harassment. This post-processing layer blocks or sanitises harmful content before it reaches customers, satisfying the stem's requirement to reduce toxic outputs while leaving the underlying model's weights and inference performance untouched.
- ✗
Provide 50 few-shot examples of safe emails in every prompt.
Why it's wrong here
Fifty few-shot examples consume substantial context on every call and still rely on pattern imitation rather than an enforced safety mechanism, so harmful outputs remain possible. Few-shot prompting suits steering format, tone or task structure, making it attractive here, but dedicated safety classifiers or moderation layers address toxicity directly.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.