Courseiva
hardMultiple Choice

PDE Practice Question: You have two versions of a classification model…

You have two versions of a classification model (v1 and v2) deployed on a Vertex AI Endpoint. You want to gradually roll out v2 to 10% of traffic, monitor performance, and if metrics are better, increase traffic to 100%. You have set up model monitoring for skew and drift. Which configuration should you use?

⚠ Common exam trap

It's easy for candidates to confuse infrastructure-level load balancing (Option B) with Vertex AI's built-in traffic splitting, or think that replica counts (Option C) control traffic distribution, when in fact traffic_split is the only parameter that directly controls request routing percentages.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Vertex AI Endpoint 'traffic_split' parameter to assign 10% of traffic to v2 and 90% to v1.

The Vertex AI Endpoint 'traffic_split' parameter allows you to direct a percentage of inference requests to different model versions deployed on the same endpoint. Setting 10% to v2 and 90% to v1 enables a gradual rollout while monitoring skew and drift, and you can adjust the split as needed. This is the native, supported method for canary deployments in Vertex AI, avoiding the complexity and latency of external load balancers.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the Vertex AI Endpoint 'traffic_split' parameter to assign 10% of traffic to v2 and 90% to v1.

    Why this is correct

    The traffic_split parameter on a Vertex AI Endpoint distributes prediction requests across deployed model versions by percentage. Setting 10% to v2 and 90% to v1 enables gradual canary rollout, letting monitoring metrics inform whether to increase v2 traffic to 100%.

  • ✗

    Deploy v2 to a separate endpoint and use a load balancer to route 10% of traffic.

    Why it's wrong here

    A separate endpoint with a load balancer splits traffic but gives no shared metrics comparison or automated promotion between versions. This pattern fits independent models needing isolation, not gradual canary rollout with metric-based traffic shifting.

  • ✗

    Create a new deployment with v2 on the same endpoint and set the 'min_replica_count' to 1 for both versions.

    Why it's wrong here

    Setting min_replica_count to 1 only guarantees capacity; it does not split traffic between versions. Vertex AI routes by the deployed model's traffic percentage, so v2 would receive either all or none. This setting suits ensuring a model stays warm, not canary rollout.

  • ✗

    Enable Vertex AI Model Monitoring on the endpoint and set up alerting for performance drop.

    Why it's wrong here

    Monitoring and alerting detect skew, drift and performance drops but never shift traffic. The stem requires routing 10% to v2 and later 100%; that needs per-model traffic split configuration on the endpoint. Monitoring alone is correct when you only need observability, not controlled rollout.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.