hardMultiple Choice
PMLE Practice Question: A company uses Vertex AI Prediction with a custom…
A company uses Vertex AI Prediction with a custom container for a TensorFlow model. They notice that after deploying a new model version, requests still go to the old version. What is the most likely cause?
⚠ Common exam trap
Google Cloud often tests the misconception that deploying a new model version automatically replaces the old one, when in fact Vertex AI requires an explicit traffic split update to shift requests to the new version.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Traffic is not split to the new model version
In Vertex AI Prediction, when you deploy a new model version to an existing endpoint, you must explicitly allocate traffic to it. By default, the new version receives 0% traffic, so all requests continue to be served by the old version. The correct fix is to update the endpoint's traffic split, for example via the console or the `gcloud ai endpoints update` command with the `--traffic-split` flag.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The custom container is not compatible with Vertex AI
Why it's wrong here
An incompatible custom container would fail the deployment or return errors, not serve the previous version successfully; traffic splitting or a stale endpoint configuration explains old-version responses. It tempts because custom containers do impose interface requirements, and incompatibility is the correct diagnosis when a container fails to start or predict.
- ✗
The model is cached and needs cache invalidation
Why it's wrong here
Vertex AI Prediction does not cache model responses in a layer that would keep serving an old version after a new one is deployed; routing is governed by endpoint traffic allocation. It tempts because caching is a real cause of stale content in web and CDN architectures, where invalidation genuinely resolves it.
- ✓
Traffic is not split to the new model version
Why this is correct
Vertex AI Prediction routes requests according to each version's traffic split percentage. Deploying a new version does not automatically shift traffic; until the split is adjusted, the old version continues serving all requests, so requests never reach the new model.
- ✗
The new model version was not deployed to the same endpoint
Why it's wrong here
Deploying the new version to a separate endpoint leaves the original endpoint's traffic split unchanged, so requests continue reaching the old model; the version must be deployed to the same endpoint. It tempts because separate endpoints are legitimate for isolation or A/B testing, where independent routing is intended.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.