hardMultiple Select
Generative AI Leader Practice Question: Deploying a Gemini-based application and needs to…
A company is deploying a Gemini-based application and needs to ensure low latency for real-time user interactions. They also want to reduce cost. Which THREE strategies should they consider? (Select 3)
⚠ Common exam trap
A common misconception tested in this exam is that increasing output tokens or fine-tuning improves speed, when in reality these actions increase computational load or add overhead, making them counterproductive for latency and cost goals.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Gemini 1.5 Flash instead of Pro
Option A is correct because Gemini 1.5 Flash is a lighter, faster model than Gemini 1.5 Pro, delivering lower latency for real-time interactions and costing less per token, which directly addresses both the latency and cost goals. Option B is correct because caching responses to frequently repeated queries avoids redundant model invocations, cutting both latency (cache hits return instantly) and cost (fewer billed tokens). Option E is correct because transformer inference cost and latency scale with input length, so trimming the context window to only the necessary tokens reduces processing time and token charges. Option C is not appropriate because increasing max output tokens generates longer responses, raising latency and cost rather than reducing them. Option D is not appropriate because full fine-tuning is expensive and does not inherently make inference faster; latency gains typically come from smaller/distilled models or optimized serving, not full fine-tuning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Gemini 1.5 Flash instead of Pro
Why this is correct
Gemini 1.5 Flash is a smaller, distilled model offering substantially lower latency and per-token cost than Pro, directly satisfying both the real-time interaction and cost-reduction constraints. It suits high-volume, less complex tasks where Pro's deeper reasoning is unnecessary.
- ✓
Implement response caching for common queries
Why this is correct
Caching stores responses to frequently repeated queries, so subsequent identical requests return instantly without a model call. This cuts both latency for real-time interactions and inference cost, satisfying the stem's dual constraints for common query patterns.
- ✗
Increase the model's max output tokens to ensure comprehensive answers
Why it's wrong here
Raising max output tokens lengthens generated responses, increasing both latency and token-based cost — the opposite of the stated goals. It is tempting because longer outputs can improve answer completeness, so it would suit scenarios prioritising thoroughness over speed or budget.
- ✗
Use full fine-tuning to make the model faster
Why it's wrong here
Full fine-tuning updates all model weights to specialise behaviour; it does not reduce inference latency or serving cost, and training itself is expensive. It is tempting because fine-tuning can shrink prompts, so it would be the correct choice when adapting a model to a narrow domain.
- ✓
Keep the context window as short as possible by trimming input
Why this is correct
Trimming input shortens the context window, directly reducing the number of tokens processed per request. Fewer input tokens lower inference compute, which cuts both latency and cost — satisfying the real-time interaction and cost-reduction constraints simultaneously. This is the most direct lever for both goals.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.