Courseiva

PMLE Serving and Scaling Models Practice Question

Which of the following is a benefit of using Vertex AI Endpoints with autoscaling and scale-to-zero?

⚠ Common exam trap

A common misconception is that autoscaling eliminates the need for a load balancer, but in Vertex AI Endpoints, the load balancer is a separate component that remains essential for request distribution even when scaling to zero.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It reduces costs by scaling down to zero replicas when no requests are received.

Vertex AI Endpoints with autoscaling and scale-to-zero allow the number of serving replicas to dynamically adjust based on incoming traffic. When no requests are received, the endpoint can scale down to zero replicas, meaning you are not charged for idle compute resources. This directly reduces operational costs compared to maintaining a minimum number of always-on instances.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It eliminates the need for a load balancer.

    Why it's wrong here

    Vertex AI Endpoints still sit behind Google-managed load balancing, which distributes requests across replicas; autoscaling changes replica count, not the need for that layer. It would be correct if the question asked what handles traffic distribution across serving replicas.

  • ✓

    It reduces costs by scaling down to zero replicas when no requests are received.

    Why this is correct

    Scale-to-zero removes all replicas when no requests arrive, so you pay nothing during idle periods while still serving traffic when demand returns. This satisfies the cost-reduction benefit, unlike always-on endpoints that bill for idle capacity.

  • ✗

    It reduces model training time.

    Why it's wrong here

    Scale-to-zero governs serving infrastructure, spinning replicas down when idle and up on demand; it has no bearing on the training pipeline that produces the model. It would be the right benefit if the question concerned reducing compute during model development.

  • ✗

    It automatically upgrades the model version.

    Why it's wrong here

    Autoscaling and scale-to-zero adjust replica counts based on traffic, leaving model artefacts and versions untouched; version upgrades come from deploying a new model to the endpoint. It would be the right benefit if the question asked about managed model rollout rather than capacity.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.