Courseiva
mediumMultiple Choice

PDE Practice Question: A team notices that the latency for online…

A team notices that the latency for online predictions from a Vertex AI endpoint has increased significantly over the past hour. The model is a large TensorFlow model deployed with automatic scaling (minReplicaCount=2, maxReplicaCount=10). The CPU utilization of the deployed instances is consistently above 85%. What is the most likely cause of the increased latency?

⚠ Common exam trap

Google Cloud often tests the distinction between symptoms of CPU saturation (queuing/latency) versus memory or GPU issues; the trap here is that candidates may incorrectly attribute latency to network or hardware driver problems when the clear indicator is sustained high CPU utilization on existing instances.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The model is CPU-bound, and the current replicas are saturated, causing queuing.

The consistently high CPU utilization (above 85%) indicates that the existing replicas are saturated, unable to process incoming requests quickly enough. When all replicas are busy, new requests are queued, which directly increases latency. Automatic scaling can add more replicas up to maxReplicaCount=10, but if the scaling is slow or the traffic spike is sudden, queuing occurs first, causing the observed latency increase.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The network latency between the client and the endpoint has increased due to regional issues.

    Why it's wrong here

    Regional network latency would affect all endpoints uniformly and would not correlate with the endpoint's own CPU utilisation. It is tempting when latency spikes suddenly, but the 85% CPU reading localises the cause to compute saturation on the deployed instances.

  • ✗

    The model is deployed with GPU acceleration, but the instances are using incorrect CUDA drivers.

    Why it's wrong here

    Incorrect CUDA drivers would cause outright inference failures or GPU initialisation errors, not a gradual latency rise with CPU above 85%. It is tempting when GPU acceleration is assumed, but the CPU saturation points to insufficient replica capacity, not driver mismatch.

  • ✗

    The model is too large for the instance memory, causing disk swapping.

    Why it's wrong here

    Disk swapping would show memory pressure and paging, not sustained CPU above 85% across replicas. It is tempting for large models, but the symptom indicates CPU-bound inference, so adding replicas or a larger machine type addresses the actual bottleneck.

  • ✓

    The model is CPU-bound, and the current replicas are saturated, causing queuing.

    Why this is correct

    Sustained CPU above 85% across replicas indicates the model is compute-bound; requests queue behind in-flight inferences, inflating latency. Automatic scaling adds replicas but each new instance still saturates, so the bottleneck is CPU capacity, not network or model size.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.