PMLE Serving and Scaling Models Practice Question
A financial services company deploys a fraud detection model on a Vertex AI Endpoint. The model must process each transaction in under 50 ms. The team notices that p99 latency spikes to 200 ms every few minutes. Logs show that the model container performs a cold start when new replicas are added, and the autoscaler frequently adds and removes replicas. The endpoint currently has minReplicaCount=1 and maxReplicaCount=10. What should they do to reduce the latency spikes while controlling cost?
⚠ Common exam trap
The trap here is focusing on increasing maximum capacity or lowering the scaling target, when the real issue is cold starts and replica thrashing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set minReplicaCount to a value that covers baseline traffic and increase the autoscaling cool-down period to avoid rapid scale-down.
Cold starts and rapid scale-down cause p99 latency spikes. Maintaining a baseline of warm replicas ensures that new requests do not hit a cold container, and extending the cool-down period prevents replicas from being removed too quickly. This stabilizes the replica count and keeps latency low without over-provisioning for peak traffic.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the model to a Vertex AI Batch Prediction job and use online predictions only for high-value transactions.
Why it's wrong here
Batch prediction introduces higher latency and is not suitable for per-transaction real-time fraud detection. The requirement is sub-50 ms per transaction, which batch processing cannot meet. This approach also changes the serving pattern and does not solve the cold-start issue for online predictions.
- ✓
Set minReplicaCount to a value that covers baseline traffic and increase the autoscaling cool-down period to avoid rapid scale-down.
Why this is correct
Keeping a baseline number of warm replicas prevents cold starts during normal traffic fluctuations. Extending the cool-down period reduces thrashing, so replicas are not removed immediately after a spike. This combination maintains low p99 latency while avoiding unnecessary replica churn, which is the main cause of the 200 ms spikes.
- ✗
Configure the endpoint to use a smaller machine type so that replicas start faster and cold starts are shorter.
Why it's wrong here
A smaller machine type may start faster, but it also has less CPU and memory, which can increase inference latency and reduce throughput. It does not eliminate cold starts or scale-down thrashing. The primary fix is to keep warm replicas and stabilize scaling behavior, not to reduce per-replica resources.
- ✗
Increase maxReplicaCount to 20 and lower the autoscaling target CPU utilization to 50%.
Why it's wrong here
Raising the maximum and lowering the target may add replicas sooner, but it does not address cold starts or scale-down thrashing. In fact, more aggressive scaling can increase the frequency of replica additions and removals, worsening the latency spikes. The root cause is the lack of warm replicas and rapid scale-down, not insufficient maximum capacity.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.