PMLE Serving and Scaling Models Practice Question
You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?
⚠ Common exam trap
Many candidates confuse multi-model serving with containerization, assuming that bundling models into a single container (Option C) is equivalent to Vertex AI's native multi-model support, but this ignores the need for traffic splitting and independent model lifecycle management.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Vertex AI Prediction with multi-model serving by deploying multiple models to one endpoint with traffic splits.
Vertex AI Prediction supports multi-model serving, allowing you to deploy multiple models to a single endpoint and use traffic splits to route a percentage of requests to each model. This reduces costs by sharing underlying infrastructure (e.g., compute resources) across models, rather than provisioning separate endpoints or containers for each model.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud Run to serve each model separately.
Why it's wrong here
Cloud Run deploys each model as an independent container, which does not consolidate multiple models onto a single endpoint; the question requires a single Vertex AI endpoint to host multiple models, and Cloud Run lacks Vertex AI’s built-in model routing and traffic-splitting capabilities. It is tempting because Cloud Run offers scalable serverless hosting for individual models, and would be correct if the goal were to isolate each model in its own endpoint for independent scaling or deployment, rather than sharing one endpoint to reduce costs.
- ✓
Use Vertex AI Prediction with multi-model serving by deploying multiple models to one endpoint with traffic splits.
Why this is correct
Multiple models can be deployed to a single endpoint, each receiving a portion of the traffic.
- ✗
Package all models into a single container and deploy that container.
Why it's wrong here
This is complex and not recommended; models should be separate for independent updates.
- ✗
Deploy each model to its own endpoint and use a load balancer.
Why it's wrong here
This increases costs and management overhead.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.