PMLE Serving and Scaling Models Practice Question
You are deploying a large deep learning model on Vertex AI endpoints. The model requires GPU acceleration and you want to minimize cold-start latency. Which TWO actions should you take? (Choose 2 correct answers)
⚠ Common exam trap
Google often tests the misconception that scale-to-zero (minReplicaCount=0) reduces latency, when in fact it increases cold-start latency; the correct approach is to keep at least one replica always warm (minReplicaCount=1) and pre-load the model during container startup.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a custom container that loads the model during startup.
Option B is correct because using a custom container that loads the model during startup lets you control and optimize the initialization process (for example, preloading weights and warming up the GPU) so the model is ready as soon as the container starts, which directly reduces cold-start latency. Option E is correct because setting minReplicaCount to 1 keeps at least one replica always provisioned and running on Vertex AI, so incoming requests hit an already-loaded model instead of triggering a new replica spin-up, eliminating the cold start entirely. Option A is wrong because setting minReplicaCount to 0 enables scale-to-zero, meaning replicas are torn down when idle and every new request incurs a full cold start, which is the opposite of the goal. Option C is wrong because increasing maxReplicaCount only raises the ceiling for horizontal scaling under load and does nothing to reduce the latency of starting an individual replica. Option D is wrong because batch prediction is an asynchronous, job-based mode that does not serve real-time online requests and therefore is irrelevant to minimizing endpoint cold-start latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set minReplicaCount to 0 to allow scale-to-zero.
Why it's wrong here
minReplicaCount of 0 permits scale-to-zero, so the first request after idle must provision a GPU node and load the model, adding cold-start latency. It is tempting because scale-to-zero cuts cost for sporadic traffic, and would be correct where minimising idle spend outweighs response time.
- ✓
Use a custom container that loads the model during startup.
Why this is correct
Loading the model during container startup means weights are already in memory when the first request arrives, rather than fetched lazily per request. This directly reduces cold-start latency for the GPU-backed deep learning model described in the stem, complementing replica-warming measures.
- ✗
Increase maxReplicaCount to a high number.
Why it's wrong here
maxReplicaCount caps horizontal scaling under load; it does not keep GPU instances warm, so an idle endpoint still scales down and incurs cold starts. It is tempting because a high ceiling improves throughput during traffic spikes, and would be correct when the goal is handling peak concurrency rather than reducing startup delay.
- ✗
Use batch prediction instead of online prediction.
Why it's wrong here
Batch prediction processes data asynchronously as a job and returns results to storage; it exposes no endpoint and cannot serve low-latency online requests. It is tempting because batch prediction is cheaper for large offline scoring volumes, and would be correct for scheduled bulk inference rather than interactive serving.
- ✓
Set minReplicaCount to 1 to always have at least one replica running.
Why this is correct
Setting minReplicaCount to 1 keeps at least one GPU-backed replica permanently provisioned, so requests never trigger a fresh model load. This eliminates cold starts entirely for the stem's deep learning endpoint, trading idle cost for consistently low first-request latency.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.