easyMultiple Choice
PMLE Practice Question: A company deploys a model on Vertex AI Endpoints…
A company deploys a model on Vertex AI Endpoints for real-time inference. They notice latency spikes during peak hours. Which action is most effective to reduce latency without sacrificing accuracy?
⚠ Common exam trap
PMLE often tests whether candidates pick a static fix (bigger machine, pruning) when the scenario describes a dynamic load problem — the trap is missing that 'peak hours' implies autoscaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling based on CPU utilization
Latency spikes during peak hours indicate the endpoint is under-provisioned for concurrent load. Enabling autoscaling based on CPU utilization (or a custom metric like request count per replica) lets Vertex AI add replicas as demand rises, absorbing the spike without changing the model or sacrificing accuracy. This is the most direct and effective action for peak-hour latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable autoscaling based on CPU utilization
Why this is correct
Autoscaling adds replica capacity when CPU utilisation rises, spreading inference requests across more nodes during peak load. This reduces per-request queueing latency while the same model and precision are served, so accuracy is unchanged. It directly addresses the peak-hour latency spikes.
- ✗
Use a larger machine type
Why it's wrong here
A larger machine type raises per-instance throughput but does not reduce the queueing delay that causes peak-hour latency; scaling the endpoint's replica count (autoscaling) distributes concurrent requests instead. Larger machines suit sustained high CPU or memory demand from a single model, not bursty traffic spikes.
- ✗
Reduce model size by pruning
Why it's wrong here
Pruning alters the model's weights and structure, which can degrade accuracy, so it does not satisfy the no-sacrifice constraint. It tempts as a latency reduction technique, but the correct approach adds replicas or autoscaling to Vertex AI Endpoints to absorb peak traffic without touching the model.
- ✗
Implement client-side caching
Why it's wrong here
Client-side caching only helps repeated identical requests and does nothing for the endpoint's own compute latency during peak load. It tempts because caching reduces perceived response time, but the stem requires reducing server-side inference latency, which endpoint autoscaling or additional replicas address.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.