PMLE Serving and Scaling Models • Set 10
PMLE Serving and Scaling Models Practice Test 10 — 15 questions with explanations. Free, no signup.
You are deploying a large language model on a Vertex AI endpoint. The model is loaded from a Cloud Storage bucket at container startup, which adds 3 minutes to each cold start. You want to reduce cold-start time and ensure predictable latency during scale-out. Which approach should you take?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.