Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?
Keeping warm replicas sized for the burst removes the cold-start gap entirely, and lowering the autoscaling metric target makes the autoscaler add replicas before saturation rather than after. Together they absorb the sudden spike immediately while still allowing scale-down when traffic subsides, which is exactly what is needed here.
Why this answer
Bursts that outpace autoscaling reaction time must be met with pre-provisioned warm capacity. Setting minReplicaCount near peak demand guarantees replicas are already serving when the spike lands, and a lower metric target causes additional replicas to be requested earlier in the ramp. Merely changing the maximum, logging, or duplicating endpoints does not remove the delay between demand and available capacity.
Exam trap
The trap here is believing that raising maxReplicaCount makes autoscaling react fast enough to absorb a sudden burst.