PMLE Serving and Scaling Models Practice Question
A retail company serves a product-ranking model on a Vertex AI endpoint. Traffic is highly predictable: a steady baseline all day with a sharp peak every evening. During the evening peak, prediction latency exceeds the SLO for several minutes before autoscaling stabilises. The team wants to reduce this scale-up lag without over-provisioning hardware for the entire day. Which configuration should they apply to the deployed model?
⚠ Common exam trap
The trap here is assuming that a larger maximum replica count automatically reduces scale-up latency, when the delay actually comes from reactive scaling that only starts after load is already observed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure a dedicated autoscaling metric with a lower utilization target and define a scale-up schedule aligned to the evening peak.
The latency breach is a capacity-timing problem, not a capacity-ceiling problem, so the fix must make replicas available before demand arrives. Vertex AI Model Deployment autoscaling lets you choose the metric the autoscaler tracks and set a target utilization, and it also supports scheduled scaling. Lowering the target makes scale-out begin earlier, and a schedule aligned with the predictable evening peak pre-warms replicas, eliminating the reactive ramp-up lag while keeping off-peak cost low.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch the endpoint to a private endpoint and increase the machine type of each replica.
Why it's wrong here
Private connectivity affects network path and security, not scaling behaviour, so it cannot shorten the time to add capacity. A larger machine type raises per-replica throughput and may delay the moment autoscaling triggers, but each new replica still takes minutes to become ready. The predictable evening spike would still be served by too few replicas at its onset.
- ✗
Enable request-response logging and reduce the model's input feature count.
Why it's wrong here
Logging is an observability feature and has no effect on how quickly replicas are provisioned. Reducing feature count may lower per-request compute, but the scenario already has a stable baseline, so the existing replicas are sufficient most of the day; the problem is purely the delay in adding capacity for the known peak. This change does not target that root cause.
- ✓
Configure a dedicated autoscaling metric with a lower utilization target and define a scale-up schedule aligned to the evening peak.
Why this is correct
Vertex AI Model Deployment autoscaling supports both a target utilization metric and scheduling options. Lowering the target utilization makes the autoscaler add replicas earlier, while a schedule pre-warms capacity exactly when the predictable evening surge begins. Together they cut the reactive ramp-up delay without paying for peak capacity around the clock, which is precisely the SLO gap described.
- ✗
Set a higher maxReplicaCount and rely on the default autoscaling metrics.
Why it's wrong here
Raising the maximum replica ceiling only changes how far the deployment can grow; it does nothing to make the first few minutes faster. The lag comes from reactive scaling reacting to already-observed load, so a larger ceiling still leaves the endpoint under-provisioned at the start of each evening peak. It also increases the worst-case cost without addressing the latency breach during ramp-up.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.