Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions
A startup wants to generate concise summaries of long news articles using an LLM on Vertex AI. They prioritize low latency and cost. Which model choice is most appropriate?
⚠ Common exam trap
Candidates often assume the most accurate model (Gemini 1.5 Pro) is always the best choice, overlooking the specific business requirements for low latency and cost, which Gemini 1.5 Flash directly addresses.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Gemini 1.5 Flash, which is designed for high throughput and low cost.
Gemini 1.5 Flash is optimized for high-throughput, low-latency, and cost-efficient summarization tasks, making it the ideal choice for a startup that needs to process long news articles quickly without incurring high costs. It balances performance and economy, whereas Gemini 1.5 Pro prioritizes accuracy at higher latency and cost, and PaLM 2 Bison is less efficient for this use case.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Gemini 1.5 Pro for the highest accuracy.
Why it's wrong here
Gemini 1.5 Pro prioritises maximum reasoning quality and long-context accuracy, incurring higher latency and per-token cost than Flash-tier models, contradicting the stated priorities. It is tempting because summarisation quality matters, and Pro would be the right pick when accuracy outweighs latency and budget.
- ✗
Use PaLM 2 Bison, as it is the most economical.
Why it's wrong here
PaLM 2 Bison is an older generation with weaker cost-per-token and latency than current Gemini Flash-class models, so it fails the low-latency, low-cost priority. It is tempting because Bison was historically positioned as the economical workhorse, and would suit workloads locked to legacy Vertex AI PaLM endpoints.
- ✗
Use Vertex AI Text Embeddings, since embeddings can generate summaries.
Why it's wrong here
Text Embeddings convert text into numeric vectors for similarity search, clustering and retrieval; they cannot decode vectors back into readable prose, so no summary is produced. It is tempting because embeddings underpin RAG pipelines that feed summarisers, but the embedding model itself generates no natural-language output.
- ✓
Use Gemini 1.5 Flash, which is designed for high throughput and low cost.
Why this is correct
Gemini 1.5 Flash is optimised for high throughput and low cost per token, matching the startup's latency and budget priorities. Summarisation of long articles needs a large context window, which Flash also provides, so it outperforms heavier Pro-tier models here.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.