Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?

⚠ Common exam trap

The trap here is believing that a separate endpoint plus an external load balancer is equivalent to an endpoint traffic split, when the native feature avoids client changes and additional infrastructure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy both model versions to the same endpoint and set a traffic split percentage.

Deploying multiple models to a single endpoint and configuring a traffic split is the native Vertex AI mechanism for canary or A/B testing. The endpoint continues to expose one URL, so clients remain unchanged, and the split percentage controls how much live traffic reaches each model version. This enables safe evaluation of the new version with minimal risk.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create two separate endpoints and use a load balancer to distribute traffic between them.

    Why it's wrong here

    Using two endpoints and an external load balancer requires changes to the client or an additional component, and it does not use the native endpoint traffic split feature. The scenario explicitly asks to split traffic at the endpoint level without changing the client, so this approach adds unnecessary complexity and does not meet the requirement.

  • ✓

    Deploy both model versions to the same endpoint and set a traffic split percentage.

    Why this is correct

    Vertex AI Endpoints support deploying multiple models to the same endpoint and assigning a traffic split percentage to each deployed model. This allows a gradual rollout to a new version while keeping the client application pointed at a single endpoint URL, which matches the requirement exactly.

  • ✗

    Deploy the new model version as a separate endpoint and update the client to call both endpoints.

    Why it's wrong here

    This requires client changes and does not provide a controlled traffic split at the endpoint. The scenario states that the client application should not be changed, so modifying the client to call both endpoints violates the constraint. It also lacks the centralized traffic management that the endpoint feature provides.

  • ✗

    Use a Vertex AI batch prediction job to send a percentage of live traffic to the new model.

    Why it's wrong here

    Batch prediction jobs process stored data in bulk and are not designed for live traffic routing. They cannot intercept online requests or split them between model versions. This approach would not provide real-time evaluation of the new model under production traffic.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.