hardMultiple Choice
PMLE Practice Question: A global retailer has deployed a real-time…
A global retailer has deployed a real-time product recommendation model on Vertex AI Endpoints. The model is a large neural network that runs on a single node with 8 vCPUs and 30 GB memory. Over the past week, the p99 latency has increased from 200ms to 2 seconds, and the error rate has risen to 5%. Cloud Monitoring shows that the endpoint's CPU utilization is consistently near 100%, and memory is at 80%. The ML engineer suspects the model is too large for the node, but model size has not changed. Logs show no increase in request volume (steady at 50 QPS). There are no recent model updates. The engineer has tried to increase the node to 16 vCPUs, but latency decreased only slightly. What is the most likely root cause and the best first step to resolve it?
⚠ Common exam trap
Google Cloud often tests the misconception that latency and CPU issues are always solved by scaling up hardware, when in fact software inefficiencies in the serving stack are a frequent root cause in ML deployments.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Profile the inference code to identify inefficient operations, such as unnecessary copies or suboptimal batch processing, and optimize the model serving logic.
The p99 latency spike and high CPU utilization despite unchanged model size and request volume indicate a software bottleneck, not a hardware one. Profiling the inference code (Option A) can reveal inefficient operations like unnecessary data copies or suboptimal batch processing that degrade performance on the existing node. Since increasing vCPUs barely helped, the root cause is likely within the serving logic, not the compute capacity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Profile the inference code to identify inefficient operations, such as unnecessary copies or suboptimal batch processing, and optimize the model serving logic.
Why this is correct
Steady QPS, unchanged model size and only marginal gains from extra vCPUs point to CPU-bound serving code rather than capacity. Profiling the inference path exposes inefficient operations such as redundant tensor copies or poor batching, which optimisation resolves.
- ✗
Add more nodes to the endpoint by enabling autoscaling to distribute the load.
Why it's wrong here
Autoscaling adds replicas to distribute request load, but request volume is steady at 50 QPS, so extra nodes cannot relieve a per-request bottleneck. Horizontal scaling is correct when QPS exceeds single-node capacity, which the logs explicitly rule out here.
- ✗
Retrain the model with a smaller architecture to reduce inference time.
Why it's wrong here
Retraining with a smaller architecture is premature; CPU saturation at unchanged model size and request rate points to inefficient inference or resource contention, not model size. Retraining is used when accuracy or model complexity genuinely requires a new architecture.
- ✗
Move the model to a machine type with more CPU cores and a GPU to accelerate inference.
Why it's wrong here
Adding a GPU does not address the actual bottleneck: the 16-vCPU test barely helped, indicating the constraint lies elsewhere, such as thread contention or memory bandwidth, not raw compute. GPUs suit throughput-heavy batch or large-matrix inference, not this saturated single-node CPU scenario.
Visual reference
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.