PMLE Serving and Scaling Models Practice Question
An engineer deploys a model to a Vertex AI endpoint with minReplicas=1 and maxReplicas=3. The endpoint receives a sudden traffic spike, but it does not scale up beyond 1 replica. The CPU utilization target is 60%. What is the most likely cause?
⚠ Common exam trap
Candidates often assume any traffic spike automatically triggers scaling, but Vertex AI's autoscaler only scales based on the configured metric (CPU utilization), not request volume directly.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The CPU utilization is below the target threshold, so the autoscaler does not add replicas.
Vertex AI's autoscaler uses CPU utilization as a metric to decide when to add replicas. If the CPU utilization remains below the 60% target threshold, the autoscaler will not trigger scale-up, even during a traffic spike. The endpoint is configured with minReplicas=1 and maxReplicas=3, but without exceeding the target, it stays at the minimum.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model is not deployed correctly.
Why it's wrong here
A failed deployment would surface as endpoint errors or zero healthy replicas, not one serving replica that simply refuses to grow. Redeployment is the fix when a model fails to load, whereas this endpoint is serving traffic and only the replica count is stuck.
- ✗
The endpoint is configured with the wrong machine type.
Why it's wrong here
Machine type determines compute and memory per replica, not whether the deployment scales; an undersized type would cause resource errors, not a stuck replica count. Machine type selection matters when a model exceeds available memory, but autoscaling failure here stems from the autoscaling metric or configuration.
- ✓
The CPU utilization is below the target threshold, so the autoscaler does not add replicas.
Why this is correct
Vertex AI autoscaling adds replicas only when average CPU utilisation exceeds the 60% target; if observed utilisation stays below that threshold, the scaler holds at minReplicas=1 regardless of request volume. The spike therefore did not drive per-replica CPU high enough to trigger scale-out, making this the likely cause.
- ✗
The endpoint is using GPU which cannot autoscale.
Why it's wrong here
GPU replicas do autoscale on Vertex AI; utilisation metrics are collected and drive scaling the same way as CPU. GPU endpoints are chosen when inference needs parallel matrix throughput, but the stem specifies a CPU utilisation target, so the accelerator type is not the blocker here.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.