PMLE Serving and Scaling Models Practice Question
You are A/B testing a new model version (challenger) against the current version (champion) on Vertex AI. You want to gradually shift traffic from champion to challenger while measuring business metrics. Which approach should you use?
⚠ Common exam trap
The trap here is assuming you need external infrastructure (load balancers, DNS, Cloud Armor) to split traffic, when Vertex AI endpoints already provide native traffic splitting as a first-class feature.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy both models to the same endpoint and use the traffic split feature to allocate percentages.
Vertex AI endpoints natively support traffic splitting between multiple deployed models on the same endpoint, allowing you to assign a percentage of prediction traffic to the champion and challenger. This is the built-in, supported mechanism for gradual rollouts and A/B experiments, and it integrates with Vertex AI's monitoring so you can measure business metrics per model version.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the challenger to a separate endpoint and use a load balancer to route a percentage of requests.
Why it's wrong here
Deploying a challenger to a separate endpoint with a load balancer routes traffic at the network layer, not at the Vertex AI model-serving layer, so it cannot use Vertex AI’s built-in traffic-splitting and monitoring for business metrics. This is tempting because a load balancer is a standard tool for gradual traffic migration in general web deployments, and it would be correct if the goal were simply to shift HTTP requests between two independent services without requiring Vertex AI’s model-evaluation pipeline.
- ✗
Use Cloud Armor to route traffic based on headers.
Why it's wrong here
Cloud Armor enforces WAF and layer-7 security policies; it cannot split requests between two Vertex AI model versions or emit per-variant prediction metrics. Vertex AI's endpoint trafficSplit performs that percentage routing. Header-based Cloud Armor rules fit canary releases of HTTP backends behind a load balancer, not model-version experiments.
- ✓
Deploy both models to the same endpoint and use the traffic split feature to allocate percentages.
Why this is correct
Deploying both models to one Vertex AI endpoint lets the traffic split feature assign a percentage to each deployed model ID, shifting requests gradually from champion to challenger. This directly satisfies the gradual traffic-shift requirement while both versions serve live predictions, so business metrics can be compared under real traffic.
- ✗
Create a new endpoint for the challenger and gradually shift DNS records.
Why it's wrong here
DNS record changes propagate slowly and are cached by resolvers, so traffic shifts cannot be controlled per-request or measured against business metrics. Vertex AI endpoints already split traffic natively via trafficSplit, giving deterministic percentage routing. DNS manipulation suits regional failover between independent deployments, not gradual model A/B testing.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.