Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A financial services company uses a Vertex AI Endpoint to serve a credit risk model. The model must always be available, even during maintenance windows, and they need to control the exact distribution of traffic across two model versions for a gradual rollout. They also want to minimize cold-start latency. Which deployment configuration should they use?

⚠ Common exam trap

The trap here is assuming that separate endpoints or batch processing can achieve the same level of control and low latency as a single endpoint with traffic splitting and minimum replicas.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy both model versions to the same endpoint, set minReplicaCount to at least 1 for each, and use traffic splitting to route a percentage of requests to each version.

Deploying both versions to the same endpoint with dedicated replicas and traffic splitting satisfies availability, precise rollout control, and low latency. It leverages Vertex AI's native capabilities for version management and gradual traffic shifting without additional infrastructure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy a single model version with minReplicaCount=0 and maxReplicaCount=10, and use a custom prediction routine to handle both versions based on request headers.

    Why it's wrong here

    Setting minReplicaCount=0 means the endpoint can scale to zero, causing cold starts and potential unavailability during scale-up. A custom prediction routine cannot manage two separate model versions with distinct artifacts; it would require loading both models into one container, complicating traffic splitting and version control.

  • ✗

    Create two separate endpoints, one for each model version, and use a global load balancer to distribute traffic equally between them.

    Why it's wrong here

    Two separate endpoints increase management overhead and do not provide native traffic splitting within Vertex AI. A global load balancer can distribute traffic but lacks integration with model versioning and metrics. This approach also does not guarantee minimal cold-start latency unless each endpoint has dedicated replicas, which is not specified.

  • ✗

    Deploy the model to a Vertex AI Batch Prediction job and schedule it to run every hour, then use a Cloud Function to serve predictions from the latest batch results.

    Why it's wrong here

    Batch prediction is designed for offline, large-scale inference, not real-time serving. It cannot provide low-latency responses or dynamic traffic splitting. Using a Cloud Function to serve cached batch results introduces staleness and does not meet the requirement for a continuously available real-time endpoint.

  • ✓

    Deploy both model versions to the same endpoint, set minReplicaCount to at least 1 for each, and use traffic splitting to route a percentage of requests to each version.

    Why this is correct

    This configuration ensures high availability with dedicated replicas for each model version, allows precise traffic splitting for gradual rollout, and minimizes cold starts by keeping at least one replica warm. It directly addresses all requirements: availability, traffic control, and latency.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.