mediumMultiple Choice
PMLE Practice Question: A model deployed on Vertex AI Prediction is…
A model deployed on Vertex AI Prediction is returning high latency for real-time requests. The model is a small TensorFlow model. Which troubleshooting step should the team take first?
⚠ Common exam trap
Google Cloud often tests the principle of 'start with the simplest infrastructure fix before optimizing the model or container,' so candidates mistakenly jump to retraining or custom containers without first checking if the instance type and scaling settings are appropriate.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check if the machine type is too small and enable autoscaling
High latency for real-time predictions from a small TensorFlow model often indicates that the serving infrastructure is under-provisioned. Checking the machine type and enabling autoscaling directly addresses whether the instance is too small to handle the request volume, which is the most common first step in diagnosing latency issues on Vertex AI Prediction.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Retrain the model with a larger batch size
Why it's wrong here
Batch size affects throughput during training or batch prediction, not per-request latency for a deployed model. Retraining also does not address serving infrastructure. The first step is inspecting the endpoint's machine type and autoscaling. Larger batch sizes are correct for improving training throughput or batch inference efficiency.
- ✓
Check if the machine type is too small and enable autoscaling
Why this is correct
A small TensorFlow model should not be inherently slow, so insufficient machine resources or lack of scaling is the likely bottleneck. Checking the machine type and enabling autoscaling addresses CPU or memory saturation before deeper model-level investigation.
- ✗
Use a custom container with optimized runtime
Why it's wrong here
A custom container addresses dependency or framework mismatches, not latency for a small TensorFlow model already served by a prebuilt container. The first step is checking whether the endpoint uses a GPU or CPU machine type and whether traffic is hitting a cold instance. Custom containers suit non-standard runtimes or libraries.
- ✗
Enable Cloud Armor to reduce traffic
Why it's wrong here
Cloud Armor filters and blocks malicious or unwanted HTTP(S) traffic at the edge; it does not reduce legitimate request latency on Vertex AI Prediction. High latency for a small model usually stems from machine type, scaling, or cold starts. Cloud Armor is correct when mitigating DDoS or applying WAF rules to a load balancer.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.