hardMultiple Choice
PMLE Practice Question: A large e-commerce company deploys a…
A large e-commerce company deploys a recommendation model on Vertex AI with autoscaling enabled. During Black Friday, traffic spikes rapidly. The autoscaler adds new instances, but new instances take several minutes to become ready (cold start). As a result, many requests time out. What should they do to mitigate this issue?
⚠ Common exam trap
Candidates often confuse scaling metrics or instance readiness with the fundamental need for pre-provisioned capacity, leading them to choose options that adjust autoscaling behavior without eliminating the cold-start latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set a higher minimum number of instances to handle the expected peak.
Setting a higher minimum number of instances ensures that a baseline capacity is always running and ready to serve traffic. This pre-warms instances, eliminating the cold-start latency during rapid traffic spikes, such as Black Friday, because new instances do not need to initialize from scratch.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger machine type to reduce the number of instances needed.
Why it's wrong here
Larger machines host fewer instances but each still loads the model from scratch, so cold-start latency per new instance is unchanged. Larger machine types suit steady high-throughput workloads needing more memory or accelerators, not burst absorption during traffic spikes.
- ✗
Configure the autoscaler to use CPU utilization metric instead of request count.
Why it's wrong here
CPU utilisation is a lagging signal: by the time CPU rises, the spike has already arrived and cold starts still occur. Request-count scaling reacts to incoming traffic sooner, but neither metric removes the model load time; pre-warmed capacity does.
- ✗
Increase the health check grace period for new instances.
Why it's wrong here
Extending the grace period only delays health-check evaluation; it does nothing to shorten the several-minute model load, so requests still time out. Grace periods suit slow-starting containers whose readiness lags behind process start, not cold-start latency that requires pre-warmed capacity.
- ✓
Set a higher minimum number of instances to handle the expected peak.
Why this is correct
Cold starts mean newly added instances cannot serve traffic for several minutes, so autoscaling alone cannot absorb Black Friday spikes. Raising the minimum instance count keeps enough warm capacity already running to handle the expected peak without waiting for scale-out.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.