PMLE Serving and Scaling Models Practice Question
Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?
⚠ Common exam trap
The trap here is believing that a separate endpoint plus an external load balancer is equivalent to an endpoint traffic split, when the native feature avoids client changes and additional infrastructure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy both model versions to the same endpoint and set a traffic split percentage.
Deploying multiple models to a single endpoint and configuring a traffic split is the native Vertex AI mechanism for canary or A/B testing. The endpoint continues to expose one URL, so clients remain unchanged, and the split percentage controls how much live traffic reaches each model version. This enables safe evaluation of the new version with minimal risk.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create two separate endpoints and use a load balancer to distribute traffic between them.
Why it's wrong here
Using two endpoints and an external load balancer requires changes to the client or an additional component, and it does not use the native endpoint traffic split feature. The scenario explicitly asks to split traffic at the endpoint level without changing the client, so this approach adds unnecessary complexity and does not meet the requirement.
- ✓
Deploy both model versions to the same endpoint and set a traffic split percentage.
Why this is correct
Vertex AI Endpoints support deploying multiple models to the same endpoint and assigning a traffic split percentage to each deployed model. This allows a gradual rollout to a new version while keeping the client application pointed at a single endpoint URL, which matches the requirement exactly.
- ✗
Deploy the new model version as a separate endpoint and update the client to call both endpoints.
Why it's wrong here
This requires client changes and does not provide a controlled traffic split at the endpoint. The scenario states that the client application should not be changed, so modifying the client to call both endpoints violates the constraint. It also lacks the centralized traffic management that the endpoint feature provides.
- ✗
Use a Vertex AI batch prediction job to send a percentage of live traffic to the new model.
Why it's wrong here
Batch prediction jobs process stored data in bulk and are not designed for live traffic routing. They cannot intercept online requests or split them between model versions. This approach would not provide real-time evaluation of the new model under production traffic.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.