hardMultiple Select
Generative AI Leader Moving a GenAI proof-of-concept to production Practice Question
A company is moving a GenAI proof-of-concept to production. They need to ensure the system can handle variable traffic and maintain low latency. Which THREE practices should they implement? (Choose 3)
⚠ Common exam trap
Google often tests the misconception that minimizing input context length universally improves performance, ignoring the trade-off with output quality, and that a single large instance is simpler and sufficient for production traffic, overlooking scalability and fault tolerance requirements.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement response caching for common queries
Response caching stores the outputs of frequently requested queries, allowing the system to serve them instantly without recomputation, which drastically reduces latency for repeated requests and offloads the underlying model. Auto-scaling for the serving infrastructure lets the system dynamically add or remove capacity as traffic varies, preserving low latency during peaks while avoiding over-provisioning during lulls. Monitoring and alerting on latency metrics provides the observability needed to detect degradation early and trigger operational responses before users are affected. Together these three practices address both variable traffic and low-latency requirements in production.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Implement response caching for common queries
Why this is correct
Caching reduces latency and cost by reusing responses for identical requests.
- ✓
Enable auto-scaling for the serving infrastructure
Why this is correct
Auto-scaling adjusts resources based on demand, ensuring performance during traffic spikes.
- ✗
Reduce the input context length to the absolute minimum
Why it's wrong here
Excessively reducing context length can degrade quality; it should be optimized but not minimized arbitrarily.
- ✓
Set up monitoring and alerting on latency metrics
Why this is correct
Monitoring ensures that latency stays within SLOs and enables proactive troubleshooting.
- ✗
Use a single, large instance to handle all traffic
Why it's wrong here
A single instance may become a bottleneck and is not cost-efficient for variable traffic.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.