Courseiva
mediumMultiple Choice

PDE Practice Question: Your organization deploys multiple versions of…

Your organization deploys multiple versions of the same model to Vertex AI Endpoint for A/B testing. You have a production model (v1) serving 90% of traffic and a candidate model (v2) serving 10%. After one week, you observe that v2 has a slightly lower AUC but significantly higher business metrics like click-through rate. The product team wants to gradually increase v2's traffic. However, you need to ensure that the overall prediction latency remains under 200 ms. Currently, the endpoint has 10 replicas for v1 and 2 replicas for v2. What is the best approach to roll out v2 while maintaining latency SLO?

⚠ Common exam trap

PDE often tests the trade-off between rapid rollout and maintaining SLOs; candidates may choose immediate 100% traffic or separate endpoints without considering autoscaling and gradual increase.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase v2's traffic split by 10% each day while also adding replicas for v2 based on CPU utilization.

The best approach is to gradually increase v2's traffic while dynamically scaling replicas based on CPU utilization. This allows controlled rollout, monitoring of latency, and ensures sufficient resources to maintain the latency SLO. Adding replicas based on CPU utilization helps handle increased load as traffic to v2 grows, preventing latency degradation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Merge v2's model into v1 by retraining v1 with v2's architecture and deploy as a single model.

    Why it's wrong here

    Retraining v1 with v2's architecture produces one model, eliminating the A/B split the product team needs to keep measuring. It is tempting as a consolidation step, and it would be right when v2 is fully validated and you want a single production artefact rather than two deployed versions.

  • ✗

    Immediately set v2 to serve 100% traffic and monitor latency; if it exceeds 200 ms, roll back.

    Why it's wrong here

    Shifting all traffic to v2 instantly removes the gradual ramp and risks breaching the 200 ms SLO, since v2 has only 2 replicas against v1's 10. It is tempting as a fast rollback-protected move, and it would fit a low-risk change where latency headroom is already proven.

  • ✓

    Increase v2's traffic split by 10% each day while also adding replicas for v2 based on CPU utilization.

    Why this is correct

    Raising the split incrementally limits risk, while autoscaling v2 replicas on CPU utilisation preserves the 200 ms latency SLO as v2's share grows. Static replica counts would let v2 traffic overwhelm its two replicas, breaching the latency constraint the stem imposes.

  • ✗

    Use a separate endpoint for v2 and route traffic at the load balancer level.

    Why it's wrong here

    A separate endpoint with load-balancer routing splits traffic outside Vertex AI's managed split, so the 90/10 weighting and shared latency SLO are no longer enforced by the endpoint. It is tempting for isolation, and it would suit independent scaling or separate monitoring of two unrelated models.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.