Courseiva

PMLE Serving and Scaling Models Practice Question

Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?

⚠ Common exam trap

The trap here is believing that raising maxReplicaCount makes autoscaling react fast enough to absorb a sudden burst.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set a higher minReplicaCount so the endpoint always keeps enough warm capacity for peak traffic, and combine it with a lower autoscaling metric target to trigger scaling earlier.

Bursts that outpace autoscaling reaction time must be met with pre-provisioned warm capacity. Setting minReplicaCount near peak demand guarantees replicas are already serving when the spike lands, and a lower metric target causes additional replicas to be requested earlier in the ramp. Merely changing the maximum, logging, or duplicating endpoints does not remove the delay between demand and available capacity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Set a higher minReplicaCount so the endpoint always keeps enough warm capacity for peak traffic, and combine it with a lower autoscaling metric target to trigger scaling earlier.

    Why this is correct

    Keeping warm replicas sized for the burst removes the cold-start gap entirely, and lowering the autoscaling metric target makes the autoscaler add replicas before saturation rather than after. Together they absorb the sudden spike immediately while still allowing scale-down when traffic subsides, which is exactly what is needed here.

  • ✗

    Increase the endpoint's maxReplicaCount and rely on the autoscaler's default metrics to add capacity as quickly as possible.

    Why it's wrong here

    Raising the ceiling alone does not help when scaling reacts only after utilization is observed; the autoscaler still needs several minutes to observe, decide, and start replicas. The burst is over before that loop completes, so requests continue to receive 429 responses during the ramp-up window.

  • ✗

    Enable request logging on the endpoint so Cloud Logging captures the 429 responses and the autoscaler can use log volume as an additional scaling signal.

    Why it's wrong here

    Request logging is observability, not a scaling input; Vertex AI autoscaling does not scale on log volume. Turning it on adds latency and logging cost while leaving the burst unprotected, so timeouts and 429 responses continue exactly as before with no capacity improvement.

  • ✗

    Deploy the model to a second endpoint and split traffic between the two endpoints so each handles roughly half of the incoming requests.

    Why it's wrong here

    Two endpoints each sized for normal traffic still lack warm capacity, and each autoscales independently with the same reaction delay. The burst is divided but not absorbed, so both endpoints throttle, and you now pay for duplicate infrastructure without solving the timeout problem.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.