mediumMultiple Select
Generative AI Leader Practice Question: Build a GenAI application that summarizes…
A company wants to build a GenAI application that summarizes customer reviews into actionable insights. The solution must be cost-effective and handle variable traffic. Which TWO strategies should they implement? (Choose 2)
⚠ Common exam trap
The trap here is equating 'best accuracy' with 'largest model' and 'lowest latency' with 'dedicated GPU' — the exam wants cost-effectiveness and variable-traffic handling, which point to small fine-tuned models and batching.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a small, fine-tuned model for summarization
Option C is correct because a small, fine-tuned foundation model delivers adequate summarization quality for a narrow task like review summarization while consuming far fewer compute resources and costing less per inference than a large general-purpose model. Option D is correct because batching multiple review-summarization requests into a single inference call amortizes model loading and per-request overhead, improving throughput and cost-efficiency, which is essential for handling variable traffic. Option A is not appropriate because a dedicated GPU cluster is expensive and provides fixed capacity that does not scale economically with variable traffic. Option B is not appropriate because using the largest available foundation model increases cost and latency without guaranteeing better results for a simple summarization task. Option E is not appropriate because stuffing all raw reviews into the prompt context inflates token usage, raises cost, and can exceed context limits without improving summarization quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy on a dedicated GPU cluster to minimize latency
Why it's wrong here
A dedicated GPU cluster bills continuously whether or not requests arrive, so idle periods during variable traffic still incur full compute cost. Serverless or pay-per-token inference scales to zero between bursts. Dedicated clusters are tempting for steady, latency-sensitive production workloads where consistent throughput justifies reserved capacity, which this scenario does not describe.
- ✗
Use the largest available foundation model for best accuracy
Why it's wrong here
A larger foundation model raises per-token inference cost and latency without improving summarisation enough to justify it, so variable traffic becomes expensive to serve. Selecting a smaller model matched to the task, with autoscaling, controls cost. Large models are tempting for maximum accuracy on complex reasoning, but review summarisation is a narrow task.
- ✓
Use a small, fine-tuned model for summarization
Why this is correct
A small, fine-tuned model delivers adequate summarisation quality at far lower inference cost than a large general-purpose model, directly satisfying the cost-effectiveness constraint. Fine-tuning also reduces prompt length by embedding task behaviour in the weights, cutting token usage per request — valuable when traffic varies and per-call spend must stay controlled.
- ✓
Implement batching of review summaries to reduce per-request overhead
Why this is correct
Batching consolidates multiple review-summarisation requests into single model calls, amortising token overhead and compute across reviews. This directly satisfies the cost-effectiveness constraint while smoothing variable traffic, since queued batches absorb bursts without provisioning for peak load. It reduces per-request inference cost without sacrificing summary quality.
- ✗
Store all raw reviews in the prompt context to avoid retrieval
Why it's wrong here
Pasting every raw review into the prompt inflates input tokens on each call, raising cost and eventually exceeding the context window as review volume grows. Retrieval or pre-filtering sends only relevant excerpts. Full-context prompting is tempting for small, static datasets where retrieval infrastructure is unnecessary overhead, but variable traffic makes token cost scale linearly.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.