Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A company wants to build a customer support chatbot that answers based on internal documentation. They use Vertex AI Search and want to ensure the model only uses retrieved documents. What should they do?
⚠ Common exam trap
The exam often tests the distinction between controlling model behavior (temperature, token limits) and controlling the source of information (grounding), leading candidates to mistakenly choose temperature or token adjustments as a solution for hallucination prevention.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable grounding with Vertex AI Search
Grounding with Vertex AI Search ensures the model's responses are strictly based on the retrieved documents from the internal documentation, preventing hallucination or reliance on pre-trained knowledge. Grounding works by providing the model with a search result context that it must use as the sole source for generating answers, effectively constraining the output to the provided documents.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Fine-tune the model on the documentation
Why it's wrong here
Fine-tuning bakes documentation into model weights, so answers come from learned parameters rather than retrieved passages, defeating the requirement to cite only retrieved documents. It is tempting because fine-tuning does specialise a model on domain text, and it would suit a fixed style or format rather than changing source content.
- ✓
Enable grounding with Vertex AI Search
Why this is correct
Grounding with Vertex AI Search constrains generation to retrieved internal documents, satisfying the requirement that answers derive only from that corpus. The model cites source passages rather than relying on parametric memory, which reduces hallucination. This directly meets the stem's constraint that the chatbot use solely retrieved documentation.
- ✗
Increase max output tokens
Why it's wrong here
Max output tokens caps response length only; it does not restrict generation to retrieved passages, so the model can still answer from parametric knowledge. Raising it is tempting when answers are being truncated mid-sentence, which is the scenario where it is the right setting.
- ✗
Set temperature to 0.0
Why it's wrong here
Temperature 0.0 makes sampling deterministic but does not prevent the model from generating content absent from the retrieved documents. It is tempting because low temperature reduces creative drift, and it would be correct when the goal is consistent, repeatable phrasing rather than grounding.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.