easyMultiple Choice
PDE Practice Question: A company deploys a machine learning model on…
A company deploys a machine learning model on Vertex AI for online predictions. The model experiences intermittent spikes in traffic, causing latency increases. Which strategy should the company use to ensure consistent low latency during traffic spikes?
⚠ Common exam trap
Google Cloud often tests the misconception that manual scaling or switching to batch prediction is a valid solution for real-time latency spikes, when in fact autoscaling is the only automated, cost-effective method for handling intermittent traffic on Vertex AI endpoints.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling on the Vertex AI endpoint with appropriate min and max nodes.
Vertex AI endpoints support autoscaling, which dynamically adjusts the number of prediction nodes based on incoming traffic. By setting appropriate min and max nodes, the endpoint can scale up during traffic spikes to maintain low latency and scale down during low traffic to reduce costs. This ensures consistent performance without manual intervention.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable autoscaling on the Vertex AI endpoint with appropriate min and max nodes.
Why this is correct
Autoscaling adds or removes endpoint nodes in response to traffic, so capacity grows during spikes rather than queuing requests on fixed replicas. Setting min nodes preserves a warm baseline and max nodes caps cost, directly addressing the intermittent latency increases described.
- ✗
Manually scale the deployed model to a larger machine type during peak hours.
Why it's wrong here
Manual resizing reacts only after latency has already risen and cannot track intermittent spikes; Vertex AI autoscaling adjusts replica count automatically against load. It is tempting because a larger machine type genuinely raises per-node throughput, which helps steady high load rather than bursty traffic.
- ✗
Reduce the number of prediction nodes to minimize overhead.
Why it's wrong here
Fewer prediction nodes reduce aggregate serving capacity, worsening queueing and latency precisely when demand peaks. It is tempting because cutting nodes appears to reduce coordination overhead, which can help an over-provisioned deployment running below capacity.
- ✗
Switch to batch prediction to handle all requests asynchronously.
Why it's wrong here
Batch prediction queues requests and returns results asynchronously, so it cannot serve online predictions at all, let alone with low latency. It is tempting because batch jobs absorb traffic spikes cheaply, which suits offline scoring of accumulated data rather than interactive requests.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.