mediumMultiple Choice
PMLE Practice Question: A machine learning engineer notices that the…
A machine learning engineer notices that the online prediction latency for a custom TensorFlow model deployed on Vertex AI has increased significantly over the past week. Cloud Monitoring shows that the CPU utilization of the endpoints remains below 40%, but the number of concurrent requests has doubled. What is the most likely cause of the latency increase?
⚠ Common exam trap
Google Cloud often tests the misconception that low CPU utilization always means there is spare capacity, when in reality the bottleneck can be request queuing or thread pool exhaustion that does not raise CPU usage.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Insufficient number of replicas for autoscaling
The CPU utilization remains below 40% while concurrent requests have doubled, indicating that the existing replicas are not saturated on CPU but are bottlenecked by request queuing or thread contention. Vertex AI autoscaling scales based on CPU utilization by default; if the threshold is not crossed, new replicas are not provisioned, causing requests to queue and latency to spike. The engineer should verify the autoscaling configuration and consider scaling on request count or reducing the CPU utilization target.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Data skew causing longer inference time
Why it's wrong here
Data skew affects training or batch pipelines, altering feature distributions and model accuracy, not online inference duration per request. It is tempting because skew is a genuine production concern, but it is the correct diagnosis when prediction quality drifts while latency and throughput stay stable.
- ✗
Memory leak in the serving container
Why it's wrong here
A memory leak would typically raise memory usage and trigger restarts or OOM errors, and CPU sitting below 40% with doubled concurrency points to request queuing, not container memory. It is tempting because leaks do cause gradual latency growth, but they are diagnosed from memory metrics, not concurrency.
- ✓
Insufficient number of replicas for autoscaling
Why this is correct
Doubling concurrent requests without CPU saturation points to queueing at the replica layer, not compute exhaustion. Vertex AI autoscaling adds replicas based on utilisation thresholds, so with CPU below 40% the autoscaler never scales out, leaving too few replicas to serve the higher concurrency, which inflates prediction latency.
- ✗
Model overfitting
Why it's wrong here
Overfitting degrades prediction accuracy on unseen data, not serving latency, and it would not correlate with doubled concurrent requests. It is tempting because overfitting is a common model-quality diagnosis, but it is the right answer when accuracy metrics on new data diverge from training performance.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.