Courseiva

PMLE Serving and Scaling Models Practice Question

You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?

⚠ Common exam trap

Google often tests the misconception that traffic splitting requires separate endpoints or client-side logic, when in fact Vertex AI provides a native traffic split configuration on a single endpoint.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Update the endpoint to split traffic between the two model versions using the traffic split configuration.

Vertex AI endpoints support a built-in traffic split configuration that allows you to gradually shift traffic between model versions deployed to the same endpoint. By updating the endpoint's traffic split percentages (e.g., from 100% old / 0% new to 0% old / 100% new over 24 hours), you can achieve a smooth, controlled rollout without changing client code or managing multiple endpoints.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Vertex AI Experiments to run an A/B test between the two versions.

    Why it's wrong here

    Vertex AI Experiments tracks and compares training runs and metrics; it does not route live prediction traffic between deployed model versions. It is tempting because A/B testing implies traffic comparison, and Experiments would be correct for evaluating model metrics offline before deployment.

  • ✗

    Deploy the new version to a separate endpoint and update your client to use the new endpoint for a percentage of requests.

    Why it's wrong here

    A separate endpoint splits traffic by client-side routing rather than the endpoint's own traffic split, so the gradual 24-hour shift is not managed centrally. It is tempting because separate endpoints isolate versions, and this would be correct for fully independent deployments with distinct scaling needs.

  • ✓

    Update the endpoint to split traffic between the two model versions using the traffic split configuration.

    Why this is correct

    Vertex AI endpoints support traffic split configuration across deployed model versions, letting you assign percentage weights to each. Adjusting these weights gradually shifts requests from old to new, achieving the controlled 24-hour rollout without redeployment.

  • ✗

    Delete the old version and redeploy the new version with a different endpoint name, then update DNS.

    Why it's wrong here

    Deleting the old version and renaming the endpoint destroys the existing endpoint, so no gradual traffic split can occur; DNS changes are coarse and cannot express percentage weights. Vertex AI's traffic-split parameter on a single endpoint shifts percentages between deployed model versions, which is what gradual rollout requires.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.