Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A streaming platform uses a large generative model for personalized content suggestions. Budget constraints require minimizing inference costs without significantly degrading quality. Which approach is most effective?
⚠ Common exam trap
A common pitfall is assuming that caching frequent prompts or upgrading to higher-end accelerators reduces per-inference costs. Caching only helps with repeated queries, not unique recommendations; hardware upgrades increase fixed costs. Distillation directly reduces model size and inference compute, aligning with cost constraints.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a distilled version of the model.
Distillation trains a smaller 'student' model to mimic a larger 'teacher' model, reducing parameter count and inference latency while retaining most of the recommendation quality. This directly addresses the budget constraint by lowering compute and memory costs per inference, making it the most effective approach among the options.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the model on higher-end accelerators to save time.
Why it's wrong here
Higher-end accelerators increase throughput per instance but raise hourly cost, so inference spend rises rather than falls. They are tempting because faster hardware shortens response time, and would be correct when the goal is reducing latency or meeting throughput targets, not minimising cost under a fixed quality bar.
- ✓
Use a distilled version of the model.
Why this is correct
Distillation trains a smaller model to reproduce the large model's outputs, cutting inference compute and cost while retaining most recommendation quality. This satisfies the budget constraint without the significant quality degradation that cruder reductions would cause.
- ✗
Implement stronger safety filters to reduce output length.
Why it's wrong here
Safety filters truncate or block outputs; they do not reduce the compute performed per inference, so token generation cost stays largely unchanged while quality drops. Filters are tempting because they cut visible output length, and would be correct when the requirement is content moderation or policy compliance, not cost reduction.
- ✗
Cache frequent prompts to avoid regeneration.
Why it's wrong here
Caching frequent prompts avoids regeneration only for identical repeated inputs; personalised suggestions vary per user, so hit rates stay low and cost barely moves. Caching is tempting because it genuinely cuts repeated inference, and would be correct for static or highly repetitive prompt workloads, not per-user personalisation.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.