mediumMultiple Choice
PDE Practice Question: A company has deployed a machine learning model…
A company has deployed a machine learning model on Vertex AI Prediction that serves real-time predictions for a customer-facing application. The model was trained using a custom container and is hosted on a single endpoint with a minimum number of nodes. Recently, the team noticed that during peak traffic, prediction latency increases significantly and some requests time out. The endpoint is configured with a baseline traffic split of 100% on the current model version. Which action should the team take to reduce latency and improve reliability?
⚠ Common exam trap
The trap is that candidates often confuse external load balancing (Option B) with autoscaling, assuming that distributing requests across multiple endpoints is equivalent to adding compute capacity. However, Vertex AI Prediction endpoints are single resources that cannot be scaled horizontally by fronting them with a load balancer—they require an internal autoscaling configuration, such as setting a higher maximum node count and a CPU utilization target, to dynamically add nodes during peak demand.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure horizontal autoscaling with a higher maximum number of nodes and set a CPU utilization target.
Configuring horizontal autoscaling with a higher maximum number of nodes and a CPU utilization target allows Vertex AI Prediction to automatically add more nodes during peak traffic, distributing the inference load and reducing latency. This directly addresses the root cause—insufficient compute resources under high demand—without requiring architectural changes or sacrificing availability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Reduce the minimum number of nodes to zero to allow scale-to-zero during low traffic.
Why it's wrong here
Reducing min nodes would increase cold start latency and not help during peak traffic.
- ✗
Place a Google Cloud Load Balancer in front of the Vertex AI endpoint to distribute requests across multiple endpoints.
Why it's wrong here
Vertex AI Prediction endpoints already have built-in load balancing; an external load balancer adds complexity without benefit.
- ✓
Configure horizontal autoscaling with a higher maximum number of nodes and set a CPU utilization target.
Why this is correct
Autoscaling allows the endpoint to add nodes during high traffic, reducing latency and preventing timeouts.
- ✗
Implement A/B testing by splitting traffic between two model versions to distribute load.
Why it's wrong here
A/B testing is for evaluating model performance, not for scaling to handle traffic spikes.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
3 more ways this is tested on PDE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company deploys a machine learning model on Vertex AI for online predictions. The model experiences intermittent spikes in traffic, causing latency increases. Which strategy should the company use to ensure consistent low latency during traffic spikes?
easy- ✓ A.Enable autoscaling on the Vertex AI endpoint with appropriate min and max nodes.
- B.Manually scale the deployed model to a larger machine type during peak hours.
- C.Reduce the number of prediction nodes to minimize overhead.
- D.Switch to batch prediction to handle all requests asynchronously.
Why A: Vertex AI endpoints support autoscaling, which dynamically adjusts the number of prediction nodes based on incoming traffic. By setting appropriate min and max nodes, the endpoint can scale up during traffic spikes to maintain low latency and scale down during low traffic to reduce costs. This ensures consistent performance without manual intervention.
Variation 2. A company deploys a machine learning model to Vertex AI for real-time predictions. After deployment, they notice that prediction latency spikes during peak traffic hours. Which approach should they take to reduce latency without sacrificing accuracy?
medium- ✓ A.Configure auto-scaling with higher min and max instances
- B.Reduce the number of input features
- C.Switch from online to batch prediction
- D.Use a larger machine type for the model
Why A: Configuring auto-scaling with higher min and max instances ensures that Vertex AI has sufficient pre-warmed replicas to handle traffic spikes without cold-start latency. This approach maintains model accuracy because it does not alter the model architecture or inference logic, only the infrastructure capacity.
Variation 3. A healthcare company deploys a model for diagnosing medical images on Vertex AI using a custom container with a TensorFlow model. The model uses a mixture of GPUs (NVIDIA T4) and CPUs. After deployment, you notice that prediction latency is highly variable: sometimes under 100ms, sometimes over 10 seconds. Investigation shows that the variability correlates with the number of concurrent requests. The endpoint has a min replicas of 1 and max replicas of 3, with target CPU utilization set to 80%. You also observe that GPU utilization remains low (<20%) even during high load. What is the most likely cause of the latency variability? A) The model is not fully utilizing GPUs due to inefficient data loading from CPU. B) The autoscaling metric (CPU utilization) is not appropriate for a GPU-bound workload; the endpoint does not scale based on GPU utilization. C) The GPU machine type is too small for the model. D) The container is not configured to use the GPU correctly.
hard- ✓ A.The autoscaling metric (CPU utilization) is not appropriate for a GPU-bound workload; the endpoint does not scale based on GPU utilization.
- B.The model is not fully utilizing GPUs due to inefficient data loading from CPU.
- C.The container is not configured to use the GPU correctly.
- D.The GPU machine type is too small for the model.
Why A: The endpoint autoscales on CPU utilization, but the workload is GPU-bound — GPU utilization stays under 20% while CPU may spike from data loading and request handling. Because the autoscaler reacts to CPU, it does not scale out when GPU saturation causes queuing, so latency balloons under concurrency. The correct fix is to autoscale on a GPU-relevant metric or use a custom metric.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.