Generative AI Leader Fundamentals of Generative AI Practice Question
A healthcare analytics team is evaluating a generative AI model for summarizing clinical notes. They observe that summaries are fluent but occasionally invent details not present in the source notes. They want a practical mitigation that reduces fabricated content without retraining the model. Which approach should they apply first?
⚠ Common exam trap
The trap here is assuming more output space or more training data fixes hallucination, when the first effective lever is constraining sampling and instructing the model to use only the supplied context.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Lower the temperature and instruct the model to answer only from the provided note text
Reducing temperature and explicitly instructing the model to rely only on the provided note constrains generation toward the source content, cutting down fabricated details without retraining. It is fast, inexpensive, and directly targets the failure mode. Longer output, higher randomness, or unrelated fine-tuning do not improve fidelity and can make hallucination worse.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Raise the temperature to encourage the model to explore alternative phrasings of the notes
Why it's wrong here
Higher temperature increases randomness and the likelihood of sampling unlikely tokens, which generally worsens hallucination rather than reducing it. Exploring alternative phrasings is irrelevant when the requirement is strict fidelity to source notes. This change moves in the opposite direction of the desired mitigation.
- ✗
Fine-tune the model on a large corpus of unrelated medical literature to improve fluency
Why it's wrong here
Fine-tuning on unrelated literature does not teach the model to stay faithful to a given note and may introduce additional unsupported associations. It also violates the constraint of avoiding retraining for this mitigation. The issue is grounding outputs in the provided text, which fine-tuning on external material does not solve.
- ✓
Lower the temperature and instruct the model to answer only from the provided note text
Why this is correct
Lowering temperature reduces random sampling, and an explicit instruction to use only the supplied note constrains the model to the given context. Together they reduce the chance of invented details while keeping the model unchanged. This is a low-cost, immediate mitigation that targets the observed hallucination without retraining.
- ✗
Increase the maximum output token limit so the model has more room to explain itself
Why it's wrong here
A larger token limit allows longer output but does not prevent fabrication; it may even give the model more space to elaborate unsupported details. The problem is content fidelity, not truncation. Extending length does not constrain the model to source material, so it fails to address the hallucination.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.