Generative AI Leader Fundamentals of Generative AI Practice Question
A healthcare company is using Vertex AI to build a generative AI assistant that helps doctors draft clinical notes. The assistant uses a fine-tuned PaLM 2 model deployed on a private endpoint. Recently, doctors have reported that the assistant takes over 30 seconds to respond, causing workflow delays. Additionally, the monthly Vertex AI costs have increased by 40% without a proportional increase in usage. The model responses are generally accurate but sometimes include irrelevant details. The company wants to improve response time and cost while maintaining acceptable quality. A review of logs shows that most requests are for similar note types (e.g., progress notes, discharge summaries) and that the same prompt is used repeatedly with minor variations. What should the company do first?
⚠ Common exam trap
A common mix-up: candidates confuse performance optimization techniques (quantization, quota increases) with the root cause of redundant requests, leading them to pick options that address symptoms rather than the fundamental pattern of repeated prompts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement response caching for common queries and batch process similar requests
The logs show that most requests are for similar note types with repeated prompts, making response caching ideal for reducing latency and cost. Caching stores responses for identical or near-identical queries, eliminating redundant inference calls, which directly addresses the 30-second response time and 40% cost increase without sacrificing quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to a larger model (e.g., Gemini 1.5 Pro) to improve response quality and reduce irrelevant details
Why it's wrong here
A larger model increases per-token latency and cost, worsening both reported problems while adding capacity for irrelevant detail. It is tempting when quality is the sole concern, and would be correct if the current model's outputs were clinically inadequate rather than accurate but slow and expensive.
- ✗
Increase the Vertex AI endpoint's maximum request quota to handle concurrent requests
Why it's wrong here
Raising the endpoint quota permits more concurrent requests but each still takes over 30 seconds and incurs the same per-call cost, so neither latency nor spend improves. It is tempting when throttling errors appear, and would be correct if requests were being rejected due to quota limits rather than served slowly.
- ✗
Apply model quantization (e.g., INT8) to reduce model size and inference time
Why it's wrong here
Quantisation shrinks weights and speeds single inferences, but it does not exploit the repeated near-identical prompts driving this latency and cost. It is tempting as a general inference optimisation, and would be correct where memory or per-token latency on a fixed model is the binding constraint rather than redundant identical requests.
- ✓
Implement response caching for common queries and batch process similar requests
Why this is correct
Caching repeated prompts and batching similar note requests cuts redundant model invocations, directly addressing the 30-second latency and 40% cost increase while preserving output quality. The logs confirm most requests share prompt patterns, making this the highest-impact first step.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.