PMLE Serving and Scaling Models Practice Question
You have deployed a model to a Vertex AI Endpoint and need to perform a canary release of a new model version to 10% of traffic. You want to monitor the new version's performance before gradually increasing its traffic share. What should you do?
⚠ Common exam trap
The trap here is thinking you need a separate endpoint or an external load balancer, when Vertex AI Endpoints already provide built-in traffic splitting.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Upload the new model as a new version and deploy it to the same endpoint with 10% traffic split.
Vertex AI Endpoints natively support deploying multiple model versions and splitting traffic by percentage. Deploying the new version to the same endpoint with a 10% split enables canary testing, and the split can be adjusted as confidence grows. Separate endpoints with external load balancing or a full cutover do not meet the controlled, gradual rollout requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the new model as a separate endpoint and use a load balancer to split traffic.
Why it's wrong here
Using a separate endpoint and an external load balancer adds architectural complexity and does not leverage Vertex AI's built-in traffic splitting. You would need to manage the load balancer configuration and health checks yourself. While it can work, it is not the recommended or simplest way to achieve a canary release on Vertex AI, and it complicates monitoring and rollback.
- ✗
Deploy the new model to the same endpoint with 100% traffic and monitor closely.
Why it's wrong here
Sending 100% of traffic to the new model is a full cutover, not a canary release. It exposes all users to potential issues and provides no controlled comparison with the previous version. The requirement is to start with 10% and gradually increase, so this approach violates the canary strategy and increases risk.
- ✗
Create a new endpoint for the new model and use Vertex AI's traffic director to shift traffic.
Why it's wrong here
Vertex AI does not provide a feature called traffic director for endpoint traffic splitting; traffic splitting is configured directly on an endpoint with multiple deployed models. Creating a separate endpoint would not automatically split traffic and would require additional components. This option invents a capability that does not exist in the described form.
- ✓
Upload the new model as a new version and deploy it to the same endpoint with 10% traffic split.
Why this is correct
Vertex AI Endpoints support multiple deployed models with configurable traffic splits. By deploying the new version to the same endpoint and assigning 10% of traffic, you can canary test it. You can then monitor metrics and gradually increase the split. This is the native, straightforward approach that integrates with Vertex AI monitoring and rollback.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.