mediumMultiple Choice
PMLE Practice Question: A team deploys a model using Vertex AI Endpoint…
A team deploys a model using Vertex AI Endpoint with automatic scaling. They observe that during traffic spikes, new instances take a long time to become ready, causing high latency for some requests. What should they configure to reduce this startup time?
⚠ Common exam trap
The trap here is conflating scaling policy knobs (max replicas, predictive autoscaling, target utilization) with startup-time reduction — only changes to the container/model itself shorten per-replica readiness time.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a custom container with a smaller footprint
A custom container with a smaller footprint reduces image pull time and container initialization overhead, which are the dominant contributors to Vertex AI Endpoint replica startup latency during scale-out. Smaller images pull faster from Artifact Registry and start faster, so new replicas become ready sooner and absorb traffic spikes with less queuing delay.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the max replicas
Why it's wrong here
Max replicas caps how many instances can run; it does not shorten the container image pull and model load that delay readiness. Raising it is tempting when demand exceeds capacity, but the stem's latency stems from startup duration, which min replica count or a smaller image addresses.
- ✓
Use a custom container with a smaller footprint
Why this is correct
Instance startup time is dominated by pulling and initialising the container image. A smaller custom container reduces image size, so new replicas become ready faster during spikes, directly addressing the slow scale-out latency described in the stem.
- ✗
Enable predictive autoscaling
Why it's wrong here
Predictive autoscaling forecasts demand and pre-warms capacity ahead of predicted peaks, but it does not reduce the time an individual instance needs to become ready. It is tempting because it addresses spikes, yet the stem asks about startup duration, which min replicas or faster image loading fixes.
- ✗
Set a higher target CPU utilization
Why it's wrong here
Target CPU utilisation governs when scaling triggers, not how fast a new instance becomes ready. Raising it is tempting to reduce churn, but it actually delays scale-out and leaves the container pull and model load time unchanged, so startup latency persists.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.