Courseiva

PMLE Serving and Scaling Models Practice Question

You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?

⚠ Common exam trap

Many candidates confuse multi-model serving with containerization, assuming that bundling models into a single container (Option C) is equivalent to Vertex AI's native multi-model support, but this ignores the need for traffic splitting and independent model lifecycle management.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Vertex AI Prediction with multi-model serving by deploying multiple models to one endpoint with traffic splits.

Vertex AI Prediction supports multi-model serving, allowing you to deploy multiple models to a single endpoint and use traffic splits to route a percentage of requests to each model. This reduces costs by sharing underlying infrastructure (e.g., compute resources) across models, rather than provisioning separate endpoints or containers for each model.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Cloud Run to serve each model separately.

    Why it's wrong here

    Cloud Run deploys each model as an independent container, which does not consolidate multiple models onto a single endpoint; the question requires a single Vertex AI endpoint to host multiple models, and Cloud Run lacks Vertex AI’s built-in model routing and traffic-splitting capabilities. It is tempting because Cloud Run offers scalable serverless hosting for individual models, and would be correct if the goal were to isolate each model in its own endpoint for independent scaling or deployment, rather than sharing one endpoint to reduce costs.

  • ✓

    Use Vertex AI Prediction with multi-model serving by deploying multiple models to one endpoint with traffic splits.

    Why this is correct

    Multiple models can be deployed to a single endpoint, each receiving a portion of the traffic.

  • ✗

    Package all models into a single container and deploy that container.

    Why it's wrong here

    This is complex and not recommended; models should be separate for independent updates.

  • ✗

    Deploy each model to its own endpoint and use a load balancer.

    Why it's wrong here

    This increases costs and management overhead.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.