hardMultiple Choice
Generative AI Leader Practice Question: Migrating a GenAI proof-of-concept to production
A company is migrating a GenAI proof-of-concept to production. During the pilot, they used a large model (e.g., Gemini 1.5 Pro) and incurred high costs. The use case is simple: generating short product descriptions from structured data. Which cost optimization strategy should they implement first?
⚠ Common exam trap
Google often tests the misconception that fine-tuning is the first step for any production optimization, when in reality, model selection and prompt engineering are cheaper and faster to implement for simple tasks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch to a smaller model like Gemini 1.5 Flash and use structured prompts
The primary cost driver in this scenario is the model size itself. Since the use case is simple (generating short product descriptions from structured data), a smaller model like Gemini 1.5 Flash can handle the task with significantly lower inference cost per token. Structured prompts further optimize by reducing token waste and ensuring consistent output, making this the most direct and impactful first step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Fine-tune a smaller model on the specific task
Why it's wrong here
Fine-tuning a smaller model would require curating a high-quality labelled dataset of product descriptions and retraining, which introduces latency and maintenance overhead for a simple task that could be solved with prompt engineering on a cheaper, smaller model. It is tempting because fine-tuning is the standard approach for adapting a model to a domain-specific task when the base model lacks the required knowledge, and would be correct if the structured data contained proprietary terminology or formatting not handled by zero-shot prompting.
- ✗
Implement batch processing to group requests
Why it's wrong here
Batch processing groups asynchronous requests, which suits non-interactive bulk workloads; short product descriptions generated on demand need immediate responses, so batching adds latency without addressing the cost driver. The primary saving comes from matching model size to task complexity, since a smaller model handles this simple structured-input task.
- ✗
Reduce the model's temperature to 0.0
Why it's wrong here
Temperature controls sampling randomness, not token pricing; setting it to 0.0 makes output deterministic but leaves per-token cost unchanged. It would matter when consistent, repeatable responses are required, yet the cost problem stems from using an oversized model for a simple task, which a smaller model resolves.
- ✓
Switch to a smaller model like Gemini 1.5 Flash and use structured prompts
Why this is correct
Gemini 1.5 Flash handles short structured generation at far lower cost per token than Pro, and structured prompts reduce token overhead further. Since the task is simple, this preserves output quality while directly addressing the high-cost constraint identified during the pilot.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.