PMLE Serving and Scaling Models Practice Question
You are using Vertex AI Prediction with a custom container that requires a large model file (5 GB). Deployment takes 10 minutes to start. You want to reduce cold start latency. Which action would be MOST effective?
⚠ Common exam trap
PMLE often tests cold-start mitigation — candidates confuse storage-level optimizations (SSD, compression) with the actual fix, which is keeping a warm replica via minReplicas.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set minReplicas to 1 to keep at least one instance always running.
Setting minReplicas to 1 keeps at least one prediction instance warm, eliminating the cold start caused by loading the 5 GB model from scratch on each new deployment. Vertex AI Prediction scales replicas based on traffic; with minReplicas=1, the model stays loaded in memory and responds immediately. Compression, local SSD, and batch prediction do not address the fundamental issue of keeping a warm online endpoint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Compress the model file and decompress on startup.
Why it's wrong here
Decompression adds CPU work at container start, so total cold-start time grows rather than shrinks; the 5 GB file still transfers from the registry. Compression suits reducing image size for storage or bandwidth-limited pulls, not latency-sensitive serving where the model must be resident before the first prediction.
- ✗
Use a machine type with local SSD to speed up model loading.
Why it's wrong here
Local SSD only accelerates reads after the image and model are already on the host; the dominant cost is pulling the 5 GB artifact from the registry over the network. Local SSD is the right lever for disk-bound training or inference I/O, not for registry download latency.
- ✗
Switch to batch prediction to avoid online cold start.
Why it's wrong here
Batch prediction removes online serving entirely, so it cannot reduce cold start latency for an online endpoint; it changes the workload type instead. It is tempting when latency is unacceptable, but the requirement is faster online startup, which needs image and model loading optimisations.
- ✓
Set minReplicas to 1 to keep at least one instance always running.
Why this is correct
Keeping minReplicas at 1 maintains a warm instance, so the 5 GB model file is already loaded and the container initialised. This directly eliminates the 10-minute cold start, since new requests hit a running replica rather than triggering a fresh container pull and model load.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. You need to deploy a model for online predictions with low latency. You want to ensure that the endpoint can handle traffic bursts without cold start. Which TWO configurations should you set? (Choose 2)
easy- A.Set maxReplicas to a high number to handle bursts.
- ✓ B.Set minReplicas to 1.
- C.Deploy the model as a custom container.
- D.Enable autoscaling with a target CPU utilisation of 30%.
- ✓ E.Use a machine type with sufficient memory for the model.
Why B: Option B (Set minReplicas to 1) is correct because keeping at least one replica always running prevents cold starts — with minReplicas=0 the endpoint scales to zero and the first request after idle time incurs model loading latency, whereas minReplicas=1 keeps a warm instance ready to serve immediately. Option E (Use a machine type with sufficient memory for the model) is correct because the model must fit entirely in memory on each replica; if the machine type has insufficient RAM, the model cannot be loaded or will swap/thrash, causing high latency and failed predictions, so adequate memory is a prerequisite for low-latency online serving. Option A is not among the marked answers: maxReplicas only caps the upper bound of scaling and does not by itself prevent cold starts or guarantee burst handling without a warm baseline. Option C is not marked: a custom container is a packaging choice and does not address latency or cold-start behavior. Option D is not marked: autoscaling on CPU at 30% is a scaling policy detail, not a guarantee against cold starts, and may even cause unnecessary scaling churn.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.