hardMultiple Choice
PMLE Practice Question: A real-time recommendation model deployed on…
A real-time recommendation model deployed on Vertex AI Endpoints is experiencing increased latency, especially during peak hours. The model is hosted on a single machine with 4 CPUs. Which set of actions should you take to diagnose and resolve the issue?
⚠ Common exam trap
Google Cloud often tests the misconception that scaling up (vertical scaling) or changing frameworks is the first step to fix latency, when the correct approach is to first diagnose capacity constraints and then scale out horizontally with autoscaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling on the endpoint and analyze request patterns to set min/max instances.
Enabling autoscaling on a Vertex AI Endpoint allows the deployment to dynamically adjust the number of serving instances based on real-time traffic, directly addressing peak-hour latency. Analyzing request patterns to set appropriate min/max instances ensures that the endpoint scales proactively without over-provisioning, which is the standard diagnostic and resolution approach for latency issues caused by insufficient capacity under variable load.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the machine type to with 32 CPUs and disable autoscaling.
Why it's wrong here
Adding CPUs addresses compute saturation but disabling autoscaling removes the replica scaling that absorbs peak-hour load, so latency persists whenever traffic spikes. It is tempting because vertical scaling is quick, and it would be right for a steadily CPU-bound model with predictable traffic where replica count need not change.
- ✗
Switch the endpoint to use GPUs and enable batch requests.
Why it's wrong here
GPUs accelerate matrix computation, but the latency stems from CPU-bound serving on one machine, and batch requests add queueing delay that worsens real-time response. It is tempting because GPUs suit large models, and it would be right for a compute-heavy deep model where throughput, not per-request latency, is the objective.
- ✓
Enable autoscaling on the endpoint and analyze request patterns to set min/max instances.
Why this is correct
Autoscaling adds replicas so peak-hour traffic spreads across more machines, removing the single 4-CPU bottleneck; analysing request patterns then sets min/max instances to match demand. This addresses the latency constraint by scaling horizontally rather than resizing one machine.
- ✗
Change the serving framework to use TensorFlow Serving with gRPC.
Why it's wrong here
TensorFlow Serving with gRPC changes the transport and serialisation layer, not the CPU capacity causing the latency, so the bottleneck remains. It is tempting because gRPC reduces payload overhead, and it would be right when the model already has sufficient compute but request marshalling or network round-trips dominate the response time.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company deploys a TensorFlow model on Vertex AI Prediction with a single node. During peak hours, inference latency increases. What should they do first to reduce latency?
easy- ✓ A.Enable autoscaling for the deployment
- B.Increase the machine type of the node
- C.Decrease the min replicas to 0
- D.Enable automatic batching of requests
Why A: Enabling autoscaling for the deployment is the correct first step because it allows Vertex AI Prediction to dynamically adjust the number of replicas based on incoming traffic. During peak hours, autoscaling can add more nodes to distribute the inference load, directly reducing latency without requiring manual intervention or over-provisioning.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.