PMLE Serving and Scaling Models • Set 4
PMLE Serving and Scaling Models Practice Test 4 — 15 questions with explanations. Free, no signup.
You are deploying a scikit-learn model to a Vertex AI endpoint using a custom container. The model artifacts are stored in a Cloud Storage bucket. During testing, predictions succeed, but you notice that each request takes several seconds because the model is loaded from Cloud Storage on every request. You want to minimize latency without changing the model. What should you do?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.