Courseiva
mediumMultiple Choice

PMLE Practice Question: A team deploys a model using Vertex AI Endpoint…

A team deploys a model using Vertex AI Endpoint with automatic scaling. They observe that during traffic spikes, new instances take a long time to become ready, causing high latency for some requests. What should they configure to reduce this startup time?

⚠ Common exam trap

The trap here is conflating scaling policy knobs (max replicas, predictive autoscaling, target utilization) with startup-time reduction — only changes to the container/model itself shorten per-replica readiness time.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a custom container with a smaller footprint

A custom container with a smaller footprint reduces image pull time and container initialization overhead, which are the dominant contributors to Vertex AI Endpoint replica startup latency during scale-out. Smaller images pull faster from Artifact Registry and start faster, so new replicas become ready sooner and absorb traffic spikes with less queuing delay.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the max replicas

    Why it's wrong here

    Max replicas caps how many instances can run; it does not shorten the container image pull and model load that delay readiness. Raising it is tempting when demand exceeds capacity, but the stem's latency stems from startup duration, which min replica count or a smaller image addresses.

  • ✓

    Use a custom container with a smaller footprint

    Why this is correct

    Instance startup time is dominated by pulling and initialising the container image. A smaller custom container reduces image size, so new replicas become ready faster during spikes, directly addressing the slow scale-out latency described in the stem.

  • ✗

    Enable predictive autoscaling

    Why it's wrong here

    Predictive autoscaling forecasts demand and pre-warms capacity ahead of predicted peaks, but it does not reduce the time an individual instance needs to become ready. It is tempting because it addresses spikes, yet the stem asks about startup duration, which min replicas or faster image loading fixes.

  • ✗

    Set a higher target CPU utilization

    Why it's wrong here

    Target CPU utilisation governs when scaling triggers, not how fast a new instance becomes ready. Raising it is tempting to reduce churn, but it actually delays scale-out and leaves the container pull and model load time unchanged, so startup latency persists.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.