A company is moving a GenAI proof-of-concept to production. They need to ensure the system can handle variable traffic and maintain low latency. Which THREE practices should they implement? (Choose 3)
Caching reduces latency and cost by reusing responses for identical requests.
Why this answer
Response caching stores the outputs of frequently requested queries, allowing the system to serve them instantly without recomputation, which drastically reduces latency for repeated requests and offloads the underlying model. Auto-scaling for the serving infrastructure lets the system dynamically add or remove capacity as traffic varies, preserving low latency during peaks while avoiding over-provisioning during lulls. Monitoring and alerting on latency metrics provides the observability needed to detect degradation early and trigger operational responses before users are affected.
Together these three practices address both variable traffic and low-latency requirements in production.
Exam trap
Google often tests the misconception that minimizing input context length universally improves performance, ignoring the trade-off with output quality, and that a single large instance is simpler and sufficient for production traffic, overlooking scalability and fault tolerance requirements.