PMLE Serving and Scaling Models Practice Question
A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?
⚠ Common exam trap
Watch out — candidates often confuse traffic splitting with scaling or load balancing, assuming that adjusting replicas or using an external load balancer is required, when Vertex AI's native `traffic_split` is the simplest and correct method for canary deployments.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the new version as a separate model on the same endpoint and use the `traffic_split` parameter in the deployment request.
Vertex AI endpoints support traffic splitting between multiple deployed models. By deploying the new model version to the same endpoint and setting `traffic_split` to 5% for the new version and 95% for the existing version, the endpoint automatically routes a corresponding proportion of inference requests to each model without any client-side changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a new endpoint for the new version and update the client to call both endpoints.
Why it's wrong here
Two separate endpoints cannot split traffic at a defined percentage; the client would need its own routing logic, and Vertex AI traffic splitting operates within one endpoint across deployed model versions. It is tempting because it isolates the new version safely, and would be correct for fully independent canary deployments with external load balancing.
- ✗
Deploy the new version and set the minimum replicas to 0, then gradually increase.
Why it's wrong here
Minimum replicas control capacity and cold-start behaviour, not the proportion of prediction requests routed to a version; setting zero simply scales the deployment down. It is tempting because it sounds like gradual rollout, and would be correct for cost control on a low-traffic endpoint rather than canary traffic splitting.
- ✗
Use Cloud Load Balancing to distribute traffic between two endpoints.
Why it's wrong here
Cloud Load Balancing distributes across backends, not across model versions inside a Vertex AI endpoint, so it cannot express a 5% split tied to deployed model IDs. It is tempting because it offers weighted traffic routing, and would be correct for splitting traffic between separate services or endpoints outside Vertex AI.
- ✓
Deploy the new version as a separate model on the same endpoint and use the `traffic_split` parameter in the deployment request.
Why this is correct
Deploying the new version as a separate model on the same Vertex AI endpoint and setting the `traffic_split` parameter to route 5% of requests to it directly satisfies the constraint of testing with a controlled fraction of live traffic before a full rollout. This mechanism uses the endpoint’s built-in traffic routing to allocate a precise percentage of inference requests to the new model version without requiring a separate endpoint or external load balancer.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.