Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A company is deploying a generative AI model for customer support. They want to reduce hallucinations while maintaining fluency. They have a large dataset of previous support conversations. Which strategy should they prioritize?
⚠ Common exam trap
Google Cloud often tests the misconception that tuning generation parameters (like temperature or beam search) can fix hallucinations, when in fact only grounding techniques like RAG or knowledge graph integration address the root cause of factual inaccuracy.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement retrieval-augmented generation (RAG) using the conversation dataset as a knowledge base.
Retrieval-augmented generation (RAG) directly addresses hallucinations by grounding the model's responses in factual, retrieved data from the conversation dataset. This approach allows the model to generate fluent, contextually relevant answers while reducing the risk of inventing information, as it retrieves actual support interactions as evidence before generating a response.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the beam search width to 10.
Why it's wrong here
Widening beam search explores more candidate sequences for higher-likelihood output, which does not add factual grounding and can even amplify confident errors. It suits translation or summarisation quality, not reducing hallucination against a conversation dataset.
- ✓
Implement retrieval-augmented generation (RAG) using the conversation dataset as a knowledge base.
Why this is correct
Retrieval-augmented generation grounds each response in passages retrieved from the conversation dataset, so answers reflect real support content rather than parametric guesses. This constrains fabrication while the underlying model preserves fluent phrasing, satisfying both the hallucination-reduction and fluency requirements.
- ✗
Fine-tune the model on the conversation dataset.
Why it's wrong here
Fine-tuning adapts style and format but does not reliably inject verifiable facts, and can still hallucinate on unseen questions. It suits tone and task specialisation; grounding answers in retrieved conversation content is what constrains fabrication while preserving fluency.
- ✗
Set the temperature to 0.1.
Why it's wrong here
Lowering temperature sharpens token selection but does not supply factual grounding, so hallucinations persist. It is the right lever for deterministic, repeatable outputs such as classification, not for injecting the support-conversation knowledge that retrieval or fine-tuning provides.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.