PMLE Serving and Scaling Models Practice Question
You deployed a model to a Vertex AI endpoint with minReplicas=0 and maxReplicas=5. After sending prediction requests, you notice the endpoint takes about 30 seconds to respond initially, but subsequent requests are fast. What is the most likely cause?
⚠ Common exam trap
Google often tests the distinction between cold start latency and persistent performance issues, so candidates may mistakenly attribute the initial delay to model size or network misconfiguration instead of recognizing the intentional scaling-to-zero behavior.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cold start occurs because the endpoint scaled down to zero.
Vertex AI endpoints with minReplicas=0 scale down to zero when idle. The first request after a period of inactivity triggers a cold start, where the endpoint must provision a new VM instance and load the model, causing a ~30-second delay. Subsequent requests are fast because the instance remains warm and handles them without provisioning overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model is too large for the machine type.
Why it's wrong here
A large model would slow every request, not just the first, so it does not explain the cold-start pattern. With minReplicas=0 the endpoint scales from zero, so the initial request waits for a replica to provision; the model size matters only when sustained latency affects all predictions.
- ✓
Cold start occurs because the endpoint scaled down to zero.
Why this is correct
With minReplicas=0, Vertex AI scales the endpoint to zero replicas when idle, so the first request must provision a container and load the model — roughly 30 seconds. Once a replica is warm, subsequent requests hit it directly, which is why later latency drops sharply.
- ✗
The VPC Service Controls are blocking the initial request.
Why it's wrong here
VPC Service Controls restrict network perimeters and would deny requests outright with a permission error, not delay one request by 30 seconds while later ones succeed. It is the right control when exfiltrating data from a protected perimeter must be prevented.
- ✗
The endpoint's autoscaling is misconfigured.
Why it's wrong here
Autoscaling governs replica count, not cold-start latency; with minReplicas=0 the first request must wait for a container to provision and load the model, which explains the 30-second delay. Autoscaling configuration would be the culprit if replicas failed to scale out under sustained load.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.