mediumMultiple Choice
PDE Practice Question: A production model deployed on Vertex AI Endpoint…
A production model deployed on Vertex AI Endpoint is experiencing high latency during traffic spikes. The current configuration uses a single replica. What is the most efficient solution?
⚠ Common exam trap
A common mistake is thinking that manually increasing the minimum replica count or using a larger machine type (static scaling) is the best way to reduce latency during spikes. However, Vertex AI Endpoint's autoscaling (with minReplicaCount=1 and maxReplicaCount=10) dynamically adjusts replicas based on traffic, avoiding over-provisioning and reducing cost. This is a key operational excellence principle in Google Cloud: design for elasticity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling with minReplicaCount=1 and maxReplicaCount=10
Enabling autoscaling with minReplicaCount=1 and maxReplicaCount=10 allows Vertex AI Endpoint to dynamically add replicas during traffic spikes, distributing inference requests across multiple instances and reducing latency. This is the most efficient solution as it scales resources up only when needed, avoiding over-provisioning and minimizing cost during low traffic periods.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set a higher min replica count (e.g., 3)
Why it's wrong here
Wastes resources during low traffic.
- ✓
Enable autoscaling with minReplicaCount=1 and maxReplicaCount=10
Why this is correct
Autoscaling adds replicas when traffic rises, distributing load across instances so per-request latency stays low, then scales back to one replica when idle. Setting minReplicaCount=1 and maxReplicaCount=10 satisfies the spike requirement efficiently without permanent over-provisioning.
- ✗
Use a larger machine type (e.g., n1-highmem-8)
Why it's wrong here
A larger machine type raises per-replica capacity but leaves one replica handling all concurrent requests, so queuing during spikes persists. Vertical scaling suits sustained high throughput on a single instance, not bursty concurrent traffic requiring horizontal distribution.
- ✗
Switch to batch prediction to handle spikes
Why it's wrong here
Batch prediction processes asynchronous jobs over stored data, not live requests, so it cannot serve the endpoint's real-time traffic at all. It is tempting because batch prediction genuinely suits large offline scoring runs where latency is irrelevant, but the scenario demands immediate per-request responses during spikes, which only autoscaling replicas provide.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.