Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A data science team needs to serve multiple versions of the same ML model on Vertex AI Endpoints for A/B testing. They want to gradually shift traffic from the current 'champion' model to a new 'challenger' model. Which feature should they use?

⚠ Common exam trap

The PMLE exam often tests the misconception that traffic splitting requires external load balancers or proxies, when in fact Vertex AI Endpoints provide this capability natively, and candidates may overlook the built-in feature in favor of more complex architectures.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy both models to the same endpoint and use traffic splitting.

Vertex AI Endpoints natively support traffic splitting, allowing you to deploy multiple model versions (e.g., champion and challenger) to the same endpoint and assign a percentage of traffic to each. This enables gradual A/B testing without additional infrastructure, as the endpoint automatically routes requests based on the configured split. Option C is correct because it leverages this built-in feature, which is designed specifically for this use case.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the challenger to a separate endpoint and use a proxy to split traffic.

    Why it's wrong here

    Deploying the challenger to a separate endpoint with a proxy splits traffic at the network layer, but Vertex AI Endpoints natively support traffic splitting across model versions within a single endpoint, which is the exact mechanism required for gradual A/B testing. This option is tempting because a proxy-based split is a common pattern for routing between independent services, and it would be correct if the team needed to compare models hosted on entirely different infrastructure or platforms, not within the same Vertex AI deployment.

  • ✗

    Use Cloud Load Balancing with weighted backend services.

    Why it's wrong here

    Cloud Load Balancing distributes traffic across backends, but Vertex AI Endpoints already provide native traffic splitting between deployed model versions, so an external load balancer adds unnecessary architecture. It would be correct for balancing generic HTTP services across VM or container backends.

  • ✓

    Deploy both models to the same endpoint and use traffic splitting.

    Why this is correct

    Vertex AI Endpoints support deploying multiple model versions to one endpoint with traffic splitting, letting you route a percentage to the challenger and gradually shift it. This satisfies the A/B testing and champion-to-challenger migration requirement.

  • ✗

    Use Vertex AI Experiments to manage model versions.

    Why it's wrong here

    Vertex AI Experiments tracks runs, parameters and metrics for comparison; it does not route prediction requests or split traffic between deployed models. It would be correct for recording and analysing training experiments, not for serving a champion/challenger split.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.