hardMultiple Choice
PDE Practice Question: You manage a team that deploys multiple versions…
You manage a team that deploys multiple versions of a computer vision model for A/B testing on Vertex AI Endpoints. You need to route a small percentage of traffic to a canary version while the rest goes to the stable version. You also need to gradually increase the canary traffic over time based on performance metrics. Which approach should you take?
⚠ Common exam trap
Google often tests the misconception that you need an external load balancer or separate endpoints for canary deployments, when in fact Vertex AI's native traffic splitting is the correct and simplest approach.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy both models to the same endpoint and configure traffic splitting percentages using the Vertex AI console or API.
Vertex AI Endpoints natively support traffic splitting between model versions deployed to the same endpoint. This allows you to assign a percentage of traffic (e.g., 5%) to a canary version and the remainder to the stable version, and then adjust the split over time via the console or API as performance metrics dictate. This approach avoids the complexity and latency of external load balancers or application-level routing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create two separate endpoints, one for each version, and use a separate load balancer to route a percentage of requests to the canary endpoint.
Why it's wrong here
Vertex AI Endpoints natively support traffic splitting across multiple deployed model versions on a single endpoint, with the percentage adjustable at any time to ramp the canary. Two endpoints behind an external load balancer cannot read Vertex's own performance metrics, so gradual metric-driven increases must be orchestrated manually.
- ✓
Deploy both models to the same endpoint and configure traffic splitting percentages using the Vertex AI console or API.
Why this is correct
Traffic splitting on one Vertex AI endpoint assigns each deployed model a percentage of prediction requests, so the canary receives a small share that you can raise incrementally. This satisfies the gradual rollout constraint without redeploying or changing the client's endpoint URL.
- ✗
Use Cloud Armor with weighted backend services to route a portion of requests to the canary version.
Why it's wrong here
Cloud Armor provides WAF and DDoS filtering at the HTTP(S) load balancer layer; its weighted backend services distribute traffic across backends, not across model versions deployed on a Vertex AI Endpoint. Traffic splitting between deployed models, adjustable as metrics arrive, is an endpoint-level capability.
- ✗
Implement feature flags in the application code to randomly select the model version for each prediction request.
Why it's wrong here
Feature flags select a model per request inside the application, bypassing Vertex AI Endpoints traffic splitting, so no endpoint-level canary percentage or metric-driven ramp is applied. It is tempting because flags are simple to deploy, and would suit gradual feature rollout rather than model traffic allocation.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.