Courseiva
mediumMultiple Choice

Canary Deployment Using Traffic Split on Vertex AI Endpoints

A machine learning team wants to deploy a new model version for canary testing, where only 5% of traffic is routed to the new version. Which Vertex AI endpoint configuration supports this?

Quick Answer

The correct answer is to configure the endpoint with a traffic split of 95% to the old version and 5% to the new version. This works because Vertex AI endpoints natively support traffic splitting between model versions, allowing you to route a precise percentage of inference requests to each deployed model without managing separate infrastructure. On the Google Professional Data Engineer exam, this tests your understanding of MLOps deployment strategies and the specific configuration options within Vertex AI—a common trap is confusing a canary deployment with creating a completely separate endpoint for testing, which defeats the purpose of gradual, controlled rollout. Remember that traffic split is a configuration property of a single endpoint, not a separate deployment. A useful memory tip: think of it as a "95/5 faucet" where you simply adjust the knob to control the flow between old and new versions, keeping the pipeline unified.

⚠ Common exam trap

Many exam-takers think canary testing requires external tools or client-side logic, but Vertex AI's built-in traffic splitting is the intended and simplest method for this purpose.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the endpoint with traffic split: 95% to old version, 5% to new version.

Vertex AI endpoints natively support traffic splitting, allowing you to route a specified percentage of requests to different model versions deployed on the same endpoint. By configuring a traffic split of 95% to the old version and 5% to the new version, you can perform canary testing without additional infrastructure or client-side logic. This is the correct and simplest approach within Vertex AI.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Have the client application randomly select which model to call with 5% probability.

    Why it's wrong here

    Client-side random selection shifts the routing decision to each application instance, so Vertex AI receives no traffic split configuration and cannot report per-model metrics or shift weights without redeploying clients. Server-side traffic splitting at the endpoint is required; client randomisation suits cases with no managed endpoint.

  • ✗

    Deploy the new version to a separate endpoint and direct 5% of users via a load balancer.

    Why it's wrong here

    A separate endpoint with external load-balancer splitting bypasses Vertex AI's native traffic split, which is configured on a single endpoint via deployed model traffic percentages. It is tempting because load balancers do distribute traffic, and would be correct if the models sat outside Vertex AI endpoints entirely.

  • ✓

    Configure the endpoint with traffic split: 95% to old version, 5% to new version.

    Why this is correct

    Vertex AI endpoints support traffic splitting by assigning percentage weights to deployed model versions. Setting 95% to the existing version and 5% to the new one routes exactly the required canary share, enabling gradual validation before full rollout.

  • ✗

    Use an A/B testing framework outside of Vertex AI to compare results.

    Why it's wrong here

    An external A/B framework splits traffic at the application layer, bypassing Vertex AI's endpoint, so the 5% split cannot be enforced server-side and metrics are not captured per deployed model. It suits offline experiment comparison, not live canary routing where the endpoint itself must divide traffic.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on PDE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company deploys a model to Vertex AI Endpoint. They want to run a canary deployment to test a new model version with 10% of traffic. How should they configure this?

medium
  • A.Deploy to a new endpoint and update the application to call both
  • B.Use Cloud Load Balancing to route traffic
  • ✓ C.Deploy the new model to the same endpoint and set traffic split
  • D.Deploy to Cloud Run and use gradual rollout

Why C: Vertex AI Endpoints natively support traffic splitting between model versions deployed to the same endpoint. By deploying the new model version to the same endpoint and setting a traffic split of 10% to the new version and 90% to the current version, the company can perform a canary deployment without changing the application code or infrastructure.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.