Courseiva
mediumMultiple Choice

PMLE Practice Question: A machine learning engineer notices that the…

A machine learning engineer notices that the online prediction latency for a custom TensorFlow model deployed on Vertex AI has increased significantly over the past week. Cloud Monitoring shows that the CPU utilization of the endpoints remains below 40%, but the number of concurrent requests has doubled. What is the most likely cause of the latency increase?

⚠ Common exam trap

Google Cloud often tests the misconception that low CPU utilization always means there is spare capacity, when in reality the bottleneck can be request queuing or thread pool exhaustion that does not raise CPU usage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Insufficient number of replicas for autoscaling

The CPU utilization remains below 40% while concurrent requests have doubled, indicating that the existing replicas are not saturated on CPU but are bottlenecked by request queuing or thread contention. Vertex AI autoscaling scales based on CPU utilization by default; if the threshold is not crossed, new replicas are not provisioned, causing requests to queue and latency to spike. The engineer should verify the autoscaling configuration and consider scaling on request count or reducing the CPU utilization target.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Data skew causing longer inference time

    Why it's wrong here

    Data skew affects training or batch pipelines, altering feature distributions and model accuracy, not online inference duration per request. It is tempting because skew is a genuine production concern, but it is the correct diagnosis when prediction quality drifts while latency and throughput stay stable.

  • ✗

    Memory leak in the serving container

    Why it's wrong here

    A memory leak would typically raise memory usage and trigger restarts or OOM errors, and CPU sitting below 40% with doubled concurrency points to request queuing, not container memory. It is tempting because leaks do cause gradual latency growth, but they are diagnosed from memory metrics, not concurrency.

  • ✓

    Insufficient number of replicas for autoscaling

    Why this is correct

    Doubling concurrent requests without CPU saturation points to queueing at the replica layer, not compute exhaustion. Vertex AI autoscaling adds replicas based on utilisation thresholds, so with CPU below 40% the autoscaler never scales out, leaving too few replicas to serve the higher concurrency, which inflates prediction latency.

  • ✗

    Model overfitting

    Why it's wrong here

    Overfitting degrades prediction accuracy on unseen data, not serving latency, and it would not correlate with doubled concurrent requests. It is tempting because overfitting is a common model-quality diagnosis, but it is the right answer when accuracy metrics on new data diverge from training performance.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.