Courseiva
hardMultiple Choice

PMLE Practice Question: A company uses Vertex AI Prediction with a custom…

A company uses Vertex AI Prediction with a custom container for a TensorFlow model. They notice that after deploying a new model version, requests still go to the old version. What is the most likely cause?

⚠ Common exam trap

Google Cloud often tests the misconception that deploying a new model version automatically replaces the old one, when in fact Vertex AI requires an explicit traffic split update to shift requests to the new version.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Traffic is not split to the new model version

In Vertex AI Prediction, when you deploy a new model version to an existing endpoint, you must explicitly allocate traffic to it. By default, the new version receives 0% traffic, so all requests continue to be served by the old version. The correct fix is to update the endpoint's traffic split, for example via the console or the `gcloud ai endpoints update` command with the `--traffic-split` flag.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The custom container is not compatible with Vertex AI

    Why it's wrong here

    An incompatible custom container would fail the deployment or return errors, not serve the previous version successfully; traffic splitting or a stale endpoint configuration explains old-version responses. It tempts because custom containers do impose interface requirements, and incompatibility is the correct diagnosis when a container fails to start or predict.

  • ✗

    The model is cached and needs cache invalidation

    Why it's wrong here

    Vertex AI Prediction does not cache model responses in a layer that would keep serving an old version after a new one is deployed; routing is governed by endpoint traffic allocation. It tempts because caching is a real cause of stale content in web and CDN architectures, where invalidation genuinely resolves it.

  • ✓

    Traffic is not split to the new model version

    Why this is correct

    Vertex AI Prediction routes requests according to each version's traffic split percentage. Deploying a new version does not automatically shift traffic; until the split is adjusted, the old version continues serving all requests, so requests never reach the new model.

  • ✗

    The new model version was not deployed to the same endpoint

    Why it's wrong here

    Deploying the new version to a separate endpoint leaves the original endpoint's traffic split unchanged, so requests continue reaching the old model; the version must be deployed to the same endpoint. It tempts because separate endpoints are legitimate for isolation or A/B testing, where independent routing is intended.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.