Generative AI Leader Fundamentals of Generative AI • Set 3
Generative AI Leader Fundamentals of Generative AI Practice Test 3 — 15 questions with explanations. Free, no signup.
You are an ML engineer at a retail company. You have deployed a generative AI model on Vertex AI to generate product descriptions. The model uses a custom container and is deployed to a single endpoint. Recently, you noticed that inference latency has increased significantly during peak hours, causing timeouts. You have checked the logs and found that the CPU utilization on the deployed instances is consistently above 90% during peak hours. The model is currently deployed with a single machine type (n1-standard-4) and no scaling. You need to reduce latency without incurring excessive cost. What should you do?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.