Courseiva

PMLE Serving and Scaling Models Practice Question

A retail company deploys a new recommendation model alongside the current champion on Vertex AI Endpoints. They want to gradually shift traffic to the challenger while monitoring business metrics (conversion rate). Which two steps are required? (Choose 2)

⚠ Common exam trap

This question tests the distinction between monitoring (which is optional after deployment) and the actual configuration steps required to shift traffic; candidates mistakenly select Cloud Monitoring (Option E) as a required step, but the question specifically asks for steps to 'gradually shift traffic,' which is accomplished by deploying to the same endpoint and setting the traffic split, not by monitoring after the fact.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy the challenger model to the same endpoint as the champion with a separate deployed model.

Option C is correct because Vertex AI Endpoints support multiple deployed models behind a single endpoint, so the challenger must be deployed as a separate DeployedModel (with its own model ID and deployed_model_id) on the same endpoint as the champion to enable side-by-side serving. Option D is correct because gradual traffic shifting is implemented by setting the endpoint's traffic_split map, which assigns percentage weights keyed by deployed_model_id (e.g., champion deployed model: 90, challenger deployed model: 10), allowing controlled canary rollout. Option A is not required because Vertex AI Experiments tracks training/experiment runs and metrics, not live endpoint traffic splitting. Option B is not required because Memorystore is a caching layer and does not perform traffic splitting or model comparison. Option E, while useful for observing conversion rate, is not a required step to shift traffic and is not marked correct; the question asks for the steps needed to perform the gradual shift itself.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Vertex AI Experiments to track the traffic split percentages.

    Why it's wrong here

    Vertex AI Experiments records and compares training runs, parameters and metrics; it does not route live prediction traffic between deployed models. Traffic splitting is configured on the Endpoint's deployed model settings. Experiments would be the right choice for tracking offline model evaluation results during development, not production rollout.

  • ✗

    Enable Cloud Memorystore to cache identical requests for both models.

    Why it's wrong here

    Memorystore caches repeated request responses to cut latency and backend load; it cannot split traffic between two deployed models or report per-version conversion. Endpoint traffic split configuration performs that rollout. Caching would be correct when serving identical high-volume queries and reducing inference cost, not for gradual champion-challenger migration.

  • ✓

    Deploy the challenger model to the same endpoint as the champion with a separate deployed model.

    Why this is correct

    Vertex AI Endpoints support multiple deployed models behind one endpoint, each with its own deployed model ID. Deploying the challenger alongside the champion on the same endpoint is the prerequisite that lets traffic be split between them.

  • ✓

    Configure traffic split in the endpoint's traffic_split field (e.g., champion:90, challenger:10).

    Why this is correct

    Setting the endpoint's traffic_split field deploys both models to one endpoint and routes a defined percentage of prediction requests to each deployed model ID, satisfying the gradual-shift constraint. This enables canary-style monitoring of conversion rate before promoting the challenger, without redeploying or interrupting the champion.

  • ✗

    Use Cloud Monitoring to track custom metrics like conversion rate per model version.

    Why it's wrong here

    Cloud Monitoring does track custom metrics, but the question asks for the two required steps to shift traffic and observe conversion per version. Traffic splitting itself is configured on the Vertex AI Endpoint, and conversion tracking needs the split in place first. Monitoring alone would be correct once per-version traffic already exists.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.