mediumMultiple Choice
PDE Practice Question: A company deploys a model to Vertex AI Endpoint
A company deploys a model to Vertex AI Endpoint. They want to run a canary deployment to test a new model version with 10% of traffic. How should they configure this?
⚠ Common exam trap
Google Cloud often tests the misconception that canary deployments require separate endpoints or external load balancers, when in fact Vertex AI Endpoints provide a built-in traffic splitting feature that handles this at the model version level.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the new model to the same endpoint and set traffic split
Vertex AI Endpoints natively support traffic splitting between model versions deployed to the same endpoint. By deploying the new model version to the same endpoint and setting a traffic split of 10% to the new version and 90% to the current version, the company can perform a canary deployment without changing the application code or infrastructure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy to a new endpoint and update the application to call both
Why it's wrong here
Two separate endpoints split traffic only if the application routes requests itself, so the 10% split is not enforced by Vertex AI. It is tempting because it isolates versions, and it would be correct for A/B testing where the client controls routing logic.
- ✗
Use Cloud Load Balancing to route traffic
Why it's wrong here
Cloud Load Balancing distributes traffic across backends such as instance groups or network endpoint groups, not across models hosted on a Vertex AI Endpoint. Vertex AI's own traffic-split configuration assigns a percentage to each deployed model ID. Cloud Load Balancing would be the right choice for splitting traffic between separate services or VM backends, not model versions within one endpoint.
- ✓
Deploy the new model to the same endpoint and set traffic split
Why this is correct
Deploying both models to one Vertex AI Endpoint and assigning a traffic split routes exactly 10% of prediction requests to the new version, satisfying the canary requirement. Vertex AI supports multiple DeployedModels per endpoint with configurable traffic percentages, so no separate endpoint or redeployment is needed.
- ✗
Deploy to Cloud Run and use gradual rollout
Why it's wrong here
Vertex AI Endpoint supports traffic splitting between deployed model versions, so a 10% canary is configured there; Cloud Run serves containerised HTTP applications and has no model-version routing. Cloud Run's gradual rollout is correct for releasing a new application revision, not for splitting inference traffic across two models.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.