hardMultiple Select
Generative AI Leader Practice Question: Deploying a document summarization solution using…
A company is deploying a document summarization solution using Vertex AI. They want to minimize cost while maintaining quality. Which three strategies should they implement? (Choose THREE)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Choose a smaller, specialized model (e.g., Gemini 1.5 Flash)
Using caching reduces repeated processing, choosing a smaller model lowers token cost, and batching requests minimizes overhead. Prompt compression is not a standard Vertex AI feature, and using the largest model increases cost.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use prompt compression to cut token usage
Why it's wrong here
Prompt compression reduces input tokens, but summarisation cost is dominated by output tokens and repeated document ingestion; compressing prompts risks dropping content the summary must retain. It suits high-volume, low-complexity classification prompts where input length drives spend, not quality-sensitive summarisation.
- ✓
Choose a smaller, specialized model (e.g., Gemini 1.5 Flash)
Why this is correct
Gemini 1.5 Flash delivers lower per-token inference pricing than larger Gemini Pro models, directly satisfying the stem's cost-minimisation constraint. Its summarisation quality remains sufficient for document summarisation, so the quality requirement holds. Selecting a smaller specialised model therefore reduces spend without materially degrading output.
- ✓
Implement response caching for repeated queries
Why this is correct
Caching responses for repeated queries avoids re-invoking the model for identical prompts, directly cutting per-token inference charges. This satisfies the cost-minimisation constraint without altering model quality, since cached outputs are reused verbatim rather than regenerated.
- ✗
Use the largest available model for best quality
Why it's wrong here
Largest models raise per-token inference cost, directly contradicting the cost-minimisation requirement while quality gains are marginal for summarisation. Selecting the largest model is right when accuracy on complex reasoning tasks outweighs budget, not when cost and quality must be balanced.
- ✓
Batch multiple document summarization requests together
Why this is correct
Batching multiple summarisation requests into a single call amortises per-request overhead and improves throughput, lowering the effective cost per document. This meets the minimise-cost constraint while preserving output quality, as each document is still summarised by the model.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.