Courseiva

PMLE Serving and Scaling Models Practice Question

You have a Vertex AI endpoint with two deployed models: a champion (v1) and a challenger (v2). You set the traffic split to 90% v1 and 10% v2. After a week, you observe that v2 has better business metrics. You want to shift all traffic to v2 gradually over 3 days to avoid any risk. What should you do?

⚠ Common exam trap

Candidates often assume deleting the old model or redeploying with 100% traffic is acceptable, but the question explicitly requires a gradual shift over 3 days to avoid risk, which only incremental traffic split updates can achieve.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Update the traffic split configuration on the endpoint multiple times over the 3 days to gradually increase v2's percentage.

Vertex AI endpoints support live traffic splitting between deployed models, allowing you to gradually shift traffic from v1 to v2 by updating the traffic split configuration multiple times over the 3-day period. This approach minimizes risk by enabling incremental rollouts and immediate rollback if issues arise, without requiring client-side changes or downtime.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy v2 to a new endpoint and update your clients to use the new endpoint.

    Why it's wrong here

    A new endpoint forces clients to change their target URL, and it provides no built-in percentage-based traffic control, so the gradual three-day shift cannot be performed. It is tempting because separate endpoints isolate deployments cleanly, but Vertex AI's existing endpoint already supports adjustable traffic splits between deployed models.

  • ✗

    Use Vertex AI Experiments to compare v1 and v2, then redeploy v2 with 100% traffic.

    Why it's wrong here

    Vertex AI Experiments tracks and compares runs; it does not control live endpoint traffic, and redeploying v2 at 100% is an immediate cutover, not the gradual three-day shift required. It is tempting because Experiments genuinely suits offline model comparison, but the stem needs endpoint traffic-split adjustments, not experiment tracking.

  • ✓

    Update the traffic split configuration on the endpoint multiple times over the 3 days to gradually increase v2's percentage.

    Why this is correct

    Updating the endpoint's traffic split repeatedly lets you raise v2's percentage incrementally, satisfying the gradual three-day shift while keeping v1 serving the remainder. Vertex AI supports modifying the deployed model traffic split on a live endpoint without redeployment, so risk is bounded at each step.

  • ✗

    Delete v1 from the endpoint so that all traffic automatically goes to v2.

    Why it's wrong here

    Deleting v1 removes the ability to split traffic, so the shift happens instantly rather than gradually over three days, defeating the stated risk-avoidance requirement. It is tempting because undeploying a model is the natural end state once v2 wins, but Vertex AI traffic splitting is the mechanism that enables a controlled, incremental rollout.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

5 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?

easy
  • A.Use Vertex AI Experiments to run an A/B test between the two versions.
  • B.Deploy the new version to a separate endpoint and update your client to use the new endpoint for a percentage of requests.
  • ✓ C.Update the endpoint to split traffic between the two model versions using the traffic split configuration.
  • D.Delete the old version and redeploy the new version with a different endpoint name, then update DNS.

Why C: Vertex AI endpoints support a built-in traffic split configuration that allows you to gradually shift traffic between model versions deployed to the same endpoint. By updating the endpoint's traffic split percentages (e.g., from 100% old / 0% new to 0% old / 100% new over 24 hours), you can achieve a smooth, controlled rollout without changing client code or managing multiple endpoints.

Variation 2. An ML engineer needs to update a model deployed on a Vertex AI endpoint without downtime. They want to gradually shift traffic to the new version while monitoring for errors. What is the correct procedure?

medium
  • A.Use a canary deployment by deploying to a separate endpoint and using a load balancer with weighted routing.
  • B.Deploy the new model to a new endpoint, then update DNS to point to the new endpoint.
  • C.Delete the old model and deploy the new one with the same endpoint.
  • ✓ D.Deploy the new model to the same endpoint with 0% traffic initially, then gradually increase traffic while monitoring.

Why D: Vertex AI supports deploying multiple models to the same endpoint and splitting traffic via a deployed model's traffic percentage. The correct canary procedure is to deploy the new model version to the existing endpoint with 0% traffic, then gradually increase its traffic share while monitoring error rates and latency. This avoids downtime and enables safe rollback by shifting traffic back to the old version.

Variation 3. You have a Vertex AI endpoint that serves a model for real-time predictions. You want to update the model to a new version with zero downtime. Which approach should you take?

easy
  • A.Delete the endpoint and recreate it with the new model.
  • ✓ B.Deploy the new model version to the same endpoint and then set traffic to 100% for the new version.
  • C.Use Cloud Load Balancing to switch traffic between two endpoints.
  • D.Create a new endpoint and update the client application to point to the new endpoint.

Why B: Vertex AI endpoints support canary deployments by allowing you to deploy a new model version to the same endpoint and then gradually shift traffic to it using the `traffic_split` parameter. Setting traffic to 100% for the new version after deployment ensures zero downtime, as the endpoint remains active and serves requests from the old version until the switch is complete.

Variation 4. A team is deploying a new model version. They want to ensure that they can quickly roll back if the new version performs poorly in production. Which TWO actions should they take? (Choose 2.)

easy
  • A.Keep the old model version deployed alongside the new one
  • B.Configure Vertex AI Model Monitoring to compare predictions
  • ✓ C.Use traffic splitting to gradually shift traffic
  • D.Set up Cloud Monitoring alerts on model performance
  • ✓ E.Store multiple model versions in the same endpoint

Why C: Option C is correct because traffic splitting on a Vertex AI endpoint lets you route only a small percentage of requests to the new model version and gradually increase it, so if the new version performs poorly you can instantly shift traffic back to the stable version. Option E is correct because deploying multiple model versions to the same endpoint is exactly what enables that traffic splitting and rollback capability, since the old version remains available to receive traffic. Options A and E overlap in intent, but A is not marked correct because simply keeping the old version deployed without an endpoint/traffic mechanism does not by itself provide controlled rollback. Option B is not correct because Model Monitoring observes skew, drift, and prediction quality; it does not perform rollback or traffic management. Option D is not correct because Cloud Monitoring alerts only notify you of performance issues, they do not enable the fast rollback action the scenario requires.

Variation 5. A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?

medium
  • A.Create a new endpoint for the new version and update the client to call both endpoints.
  • B.Deploy the new version and set the minimum replicas to 0, then gradually increase.
  • C.Use Cloud Load Balancing to distribute traffic between two endpoints.
  • ✓ D.Deploy the new version as a separate model on the same endpoint and use the `traffic_split` parameter in the deployment request.

Why D: Vertex AI endpoints support traffic splitting between multiple deployed models. By deploying the new model version to the same endpoint and setting `traffic_split` to 5% for the new version and 95% for the existing version, the endpoint automatically routes a corresponding proportion of inference requests to each model without any client-side changes.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.