easyMultiple Choice
PMLE Practice Question: A company deploys a model on Vertex AI Prediction…
A company deploys a model on Vertex AI Prediction for real-time inference. Users report intermittent high latency during peak hours. The model is deployed on a single machine type with `min_replica_count=1` and `max_replica_count=5`. Autoscaling is enabled based on CPU utilization. What is the most likely cause of the latency spikes?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Autoscaling based on CPU utilization does not react quickly to inference request spikes.
CPU utilization may not be a good proxy for inference load; the system may not scale up fast enough under sudden traffic bursts. Option A is wrong because Vertex AI automatically manages container health. Option C is wrong because Vertex AI endpoints automatically distribute traffic. Option D is wrong because the container image is built correctly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model server is crashing under load due to memory issues.
Why it's wrong here
Crashing replicas would surface as failed requests or restarts, not periodic latency, and CPU-based autoscaling already responds to load. Memory exhaustion is tempting because undersized containers do cause outages, and it would be correct if logs showed OOMKilled events or error-rate spikes rather than latency.
- ✓
Autoscaling based on CPU utilization does not react quickly to inference request spikes.
Why this is correct
CPU utilisation lags actual inference demand: it only rises after requests queue, so the autoscaler reacts after latency has already spiked. With min_replica_count=1, peak-hour bursts hit a single replica before new ones provision, producing the intermittent spikes described.
- ✗
The load balancer is misconfigured and routes traffic unevenly.
Why it's wrong here
Vertex AI Prediction manages request distribution across replicas, so uneven routing is not the operative fault; the stem's CPU-triggered autoscaling reacting slowly to peak load is. Misconfigured load balancing is tempting because it causes latency elsewhere, and it would be correct on self-managed serving infrastructure.
- ✗
The container image is not optimized for the model.
Why it's wrong here
An unoptimised image would produce consistently elevated latency from deployment onward, not intermittent spikes confined to peak hours. Image optimisation is tempting because it genuinely affects inference speed, and it would be correct if latency were uniformly poor regardless of traffic volume.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.