Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A real-time customer support chatbot using Gemini is experiencing high latency. The team must maintain response quality while improving speed. Which technique should they implement?
⚠ Common exam trap
Google often tests the misconception that latency improvements come from model parameter tuning (like temperature) or throughput adjustments (like batch size), when the real solution for real-time systems is architectural optimization like caching.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use context caching for frequent queries
Context caching reduces latency by storing frequently accessed query responses or intermediate computations, allowing the chatbot to reuse precomputed results instead of reprocessing identical or similar requests through the full model pipeline. This directly addresses high latency in real-time systems while preserving response quality, as the cached outputs are identical to freshly generated ones for the same input.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to a larger model
Why it's wrong here
A larger model increases inference latency, directly worsening the response-time problem the team must solve. Larger models are tempting because they raise output quality, and would be the right choice when accuracy is the constraint and latency budgets are generous, not when real-time speed must improve.
- ✗
Increase the batch size
Why it's wrong here
Increasing batch size raises throughput for bulk offline jobs but does not reduce per-request latency for a single interactive chatbot turn, and can even delay individual responses. Batching is tempting because it improves GPU utilisation, and would be correct for high-volume batch inference rather than real-time chat.
- ✓
Use context caching for frequent queries
Why this is correct
Context caching stores precomputed attention states for repeated prompt prefixes, so Gemini skips reprocessing them on each call. For a support chatbot with frequent recurring queries, this cuts time-to-first-token while preserving the full model and response quality, directly addressing the latency constraint without downgrading the model.
- ✗
Decrease the temperature
Why it's wrong here
Temperature controls output randomness, not the number of tokens generated or the model's compute cost, so latency is unchanged. Lowering temperature is tempting because it makes replies more deterministic, and would be correct when consistency or reduced hallucination is required, not when speed is the bottleneck.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.