PMLE Serving and Scaling Models Practice Question
A media company uses a Vertex AI endpoint to serve a video recommendation model. The model is updated weekly with new embeddings. They want to minimize downtime during model updates and ensure that the new model performs well before fully rolling it out. They also need to be able to revert quickly if issues arise. What should they do?
⚠ Common exam trap
The trap here is thinking that a new endpoint is required for testing, when Vertex AI endpoints support multiple model versions and traffic splitting natively.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the new model as a new version on the same endpoint, set it to receive 0% of traffic initially, then gradually increase traffic while monitoring performance.
The optimal approach is to deploy the new model as a version on the existing endpoint and use traffic splitting to gradually shift traffic. This provides zero-downtime updates, allows performance validation with real traffic, and enables quick rollback by adjusting traffic percentages. Other methods either cause downtime, lack gradual rollout, or add unnecessary complexity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the new model to a new endpoint and use a load balancer to distribute traffic between the old and new endpoints.
Why it's wrong here
Using a load balancer adds complexity and does not integrate with Vertex AI's built-in traffic splitting. It requires managing two endpoints and a load balancer, which can introduce latency and failure points. Vertex AI endpoints already support traffic splitting natively, making this approach redundant and less efficient.
- ✗
Create a new endpoint for the new model and update the application to call the new endpoint after testing.
Why it's wrong here
Creating a new endpoint requires application changes and does not provide a built-in mechanism for gradual traffic shifting. It may cause downtime if not handled carefully, and reverting would require another application change. This approach lacks the flexibility and safety of traffic splitting on a single endpoint.
- ✗
Use a batch prediction job to test the new model on a large dataset, then replace the existing model on the endpoint.
Why it's wrong here
Batch prediction does not test the model under real-time serving conditions and does not support gradual rollout. Replacing the model directly would cause downtime and does not allow for quick reversion. This method is unsuitable for minimizing downtime and ensuring performance in production.
- ✓
Deploy the new model as a new version on the same endpoint, set it to receive 0% of traffic initially, then gradually increase traffic while monitoring performance.
Why this is correct
Deploying a new model version on the same endpoint allows for seamless traffic splitting without downtime. Starting with 0% traffic lets you validate the new model with a small percentage of live traffic or via direct prediction requests. Gradually increasing traffic enables A/B testing and performance monitoring. If issues arise, you can quickly shift traffic back to the previous version.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.