PMLE Serving and Scaling Models Practice Question
A retail company has deployed a scikit-learn model to a Vertex AI endpoint. The model's predictions are used to personalize the homepage. During a flash sale, the endpoint experiences a sudden 10x traffic spike, and the autoscaling configuration is set to minReplicaCount=1, maxReplicaCount=3. The endpoint becomes unresponsive. You need to modify the deployment to handle similar spikes while keeping costs low during normal hours. What should you do?
⚠ Common exam trap
The trap here is assuming that CPU utilization is always the best autoscaling metric, when in fact request-based metrics can be more responsive for sudden traffic spikes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Keep minReplicaCount=1 but set maxReplicaCount to 10, and configure autoscaling based on a custom metric that tracks the number of incoming requests per second.
The best solution is to increase the maximum replica count to handle spikes and use a custom metric that directly measures request load, such as requests per second, for autoscaling. This allows rapid scaling during flash sales while keeping the minimum replicas low to control costs. CPU-based autoscaling may not react quickly enough for sudden spikes, and raising the minimum replicas increases baseline cost unnecessarily.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase maxReplicaCount to 10 and configure the autoscaling metric to CPU utilization with a target of 60%.
Why it's wrong here
Increasing max replicas helps, but CPU utilization alone may not reflect the actual load for a scikit-learn model that is memory-bound or has high per-request latency. A target of 60% might still cause throttling under sudden spikes, and the metric may lag. This does not guarantee responsiveness during flash sales and could lead to over-provisioning during normal hours if the metric is not well-tuned.
- ✗
Deploy the model to a new endpoint with minReplicaCount=1 and maxReplicaCount=10, and use a traffic split to gradually shift traffic from the old endpoint.
Why it's wrong here
Deploying a new endpoint and using traffic splitting does not address the autoscaling configuration issue. The new endpoint would still need appropriate autoscaling settings; otherwise, it would face the same problem. This adds complexity and potential downtime during the migration, and it does not inherently improve the ability to handle spikes unless the autoscaling is correctly configured.
- ✗
Set minReplicaCount to 3 and maxReplicaCount to 10, and enable autoscaling based on CPU utilization with a target of 80%.
Why it's wrong here
Raising the minimum replica count to 3 increases cost during normal hours, contradicting the requirement to keep costs low. A CPU target of 80% is too high for a latency-sensitive service; it will allow the CPU to saturate before scaling, leading to increased response times. This configuration does not optimally balance cost and performance.
- ✓
Keep minReplicaCount=1 but set maxReplicaCount to 10, and configure autoscaling based on a custom metric that tracks the number of incoming requests per second.
Why this is correct
This approach allows the endpoint to scale out rapidly during traffic spikes by using a request-based metric that directly reflects load, while keeping the minimum replica count at 1 to save costs during idle periods. A custom metric such as requests per second is more responsive for sudden spikes than CPU utilization, which may lag. Increasing max replicas provides headroom.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.