Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A data scientist is using Vertex AI generative AI studio to create a chatbot. The chatbot gives inconsistent answers to similar questions. Which parameter should they adjust to make responses more consistent?
⚠ Common exam trap
Google Cloud often tests the misconception that increasing top-p or adjusting penalties improves consistency, when in fact temperature is the primary parameter for controlling output determinism.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decrease temperature to 0.2
Decreasing the temperature to 0.2 reduces the randomness of the model's token sampling, making the output more deterministic and consistent. Temperature controls the probability distribution over tokens; lower values make the model more likely to choose the highest-probability token, reducing variability in responses to similar questions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Decrease temperature to 0.2
Why this is correct
Lowering temperature to 0.2 sharpens the probability distribution over the next token, so the model favours its highest-likelihood continuations rather than sampling broadly. That directly addresses the inconsistent answers described in the stem, since similar prompts then converge on near-identical outputs.
- ✗
Increase top-p to 0.9
Why it's wrong here
Top-p of 0.9 restricts sampling to the smallest token set whose cumulative probability reaches 90%, which still admits many low-probability tokens and yields varied outputs. Lowering top-p, not raising it, narrows sampling; high values suit diverse creative text.
- ✗
Increase presence penalty to 0.5
Why it's wrong here
Presence penalty penalises tokens already present in the output, encouraging topic diversity; raising it to 0.5 increases variation between responses rather than making them consistent. It suits creative generation where novel wording is wanted, not deterministic chatbot replies.
- ✗
Decrease frequency penalty to 0.0
Why it's wrong here
Frequency penalty penalises repeated token usage by count; setting it to 0.0 merely removes that penalty and does not constrain sampling randomness, so variance across similar prompts persists. It is used to discourage verbatim repetition within a single response.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.