mediumMultiple Select
Generative AI Leader Practice Question: A machine learning engineer wants to reduce the…
A machine learning engineer wants to reduce the latency of a Gemini-based chatbot running in production. Which TWO strategies would be MOST effective?
⚠ Common exam trap
A common misconception in Google exams is that streaming reduces total latency, but it only improves perceived latency (time-to-first-token) while total processing time remains unchanged.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch from Gemini Pro to Gemini Flash
Option A is correct because Gemini Flash is a lighter, lower-latency model variant than Gemini Pro, so switching to it directly reduces the time the model takes to generate each response. Option B is correct because max_output_tokens caps the number of tokens the model generates, and since generation time scales roughly with output length, lowering this value shortens the response and thus the overall latency. Option C is not correct because streaming improves perceived latency by delivering tokens incrementally, but it does not reduce the actual total generation time. Option D is not correct because fine-tuning adapts the model to a task and does not inherently lower inference latency. Option E is not correct because raising temperature to 1.0 affects randomness and creativity, not speed, and can even make outputs longer or less predictable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Switch from Gemini Pro to Gemini Flash
Why this is correct
Gemini Flash is a smaller, optimised model variant with substantially lower inference latency than Gemini Pro, so swapping models reduces per-request response time. This directly addresses the production chatbot's latency constraint without changing the application logic.
- ✓
Reduce the max_output_tokens parameter
Why this is correct
Capping max_output_tokens directly shortens the decode phase, since each generated token requires a separate forward pass through the model. Fewer output tokens means less sequential autoregressive computation, which is the dominant latency contributor for long responses. This satisfies the stem's production latency constraint without altering model quality or requiring infrastructure changes.
- ✗
Enable streaming mode
Why it's wrong here
Streaming returns tokens incrementally, so perceived time-to-first-token drops, but total generation latency is unchanged. It is tempting because users see output sooner, yet the stem asks for reduced latency, which streaming masks rather than removes.
- ✗
Fine-tune the model on the specific task
Why it's wrong here
Fine-tuning alters model weights for task accuracy, not inference speed; a tuned model still runs the same forward pass. It is tempting because fine-tuning reduces prompt length, but that is an indirect prompt-engineering effect, not the mechanism this scenario needs.
- ✗
Increase the temperature to 1.0
Why it's wrong here
Temperature controls output randomness, not token generation speed, so raising it to 1.0 adds variability without cutting latency. It is tempting because sampling parameters feel like tuning knobs, but temperature is chosen for creativity, not throughput.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.