mediumMultiple Choice
PMLE Practice Question: A team wants to deploy two versions of a model…
A team wants to deploy two versions of a model (v1 and v2) on Vertex AI Endpoint to conduct an A/B test. They need to split traffic so that 10% of requests go to v2. Which configuration achieves this?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy both versions on the same endpoint and use the `traffic_split` parameter to allocate 90% to v1 and 10% to v2.
Vertex AI Endpoints allow distributing traffic between deployed models using the traffic_split parameter. Setting 90% to v1 and 10% to v2 achieves the A/B test traffic split. Option B is incorrect because a global load balancer in front of two endpoints adds unnecessary complexity and is not the native Vertex AI method. Option C is incorrect because the client randomly selecting endpoints introduces client-side logic and does not leverage Vertex AI's built-in traffic splitting. Option D is incorrect because Cloud Deployment Manager's canary rollout is for infrastructure deployment, not for model traffic splitting; Vertex AI provides traffic splitting for A/B testing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Deploy both versions on the same endpoint and use the `traffic_split` parameter to allocate 90% to v1 and 10% to v2.
Why this is correct
One endpoint can host both versions as separate DeployedModels, and the traffic_split parameter maps each version's ID to a percentage. Assigning 90% to v1 and 10% to v2 delivers exactly the required A/B distribution through a single endpoint.
- ✗
Configure a global load balancer in front of two endpoints and set the weight.
Why it's wrong here
A global load balancer splits traffic across endpoints, but Vertex AI's own traffic-split configuration is bypassed, so the 10% weighting is not managed or monitored by the endpoint. Load balancers suit routing across distinct services, not version splits within one endpoint.
- ✗
Create two separate endpoints, one for each version, and have the client randomly select the endpoint.
Why it's wrong here
Two endpoints with client-side random selection moves splitting logic into application code, and the endpoint records no traffic distribution, so monitoring and adjustment are lost. Separate endpoints suit isolating models with different resources or regions, not A/B testing versions on one endpoint.
- ✗
Deploy v2 as a canary deployment and set the canary rollout to 10% in Cloud Deployment Manager.
Why it's wrong here
Cloud Deployment Manager is Google Cloud's infrastructure provisioning tool and has no canary rollout capability for Vertex AI models; the 10% setting cannot be applied there. Canary concepts belong to Cloud Run or GKE deployments, whereas Vertex AI endpoints use their own traffic-split percentages.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.