easyMultiple Choice
PDE Practice Question: A data science team needs to ensure that a…
A data science team needs to ensure that a deployed Vertex AI model can handle varying traffic patterns with minimal latency and cost. What should they do?
⚠ Common exam trap
Google Cloud often tests the misconception that batch prediction can substitute for online serving in variable traffic scenarios, but the key distinction is that batch prediction lacks real-time latency guarantees and cannot scale dynamically per request.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Vertex AI Prediction with autoscaling
Vertex AI Prediction with autoscaling dynamically adjusts the number of serving instances based on incoming traffic, ensuring minimal latency during spikes and cost efficiency during lulls. This is the recommended approach for handling variable traffic patterns in production, as it leverages Google Cloud's managed infrastructure to scale from zero to thousands of nodes automatically.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Vertex AI Prediction with autoscaling
Why this is correct
Autoscaling adjusts replica count to match incoming request volume, absorbing traffic spikes while scaling down during quiet periods. This satisfies both constraints in the stem: minimal latency under load and minimal cost when demand drops.
- ✗
Use batch prediction instead of online
Why it's wrong here
Batch prediction processes stored data asynchronously, so it cannot serve live per-request inference at low latency; it also forfeits autoscaling to demand. It is the right choice for large offline scoring jobs, not for a deployed model that must answer varying real-time traffic.
- ✗
Pre-warm all instances
Why it's wrong here
Pre-warming instances fixes capacity in advance, so it cannot track varying traffic; idle instances still bill while spikes beyond the pre-warmed count trigger cold starts. It suits predictable, scheduled peaks where demand is known ahead of time, not the autoscaling that variable load with minimal latency and cost requires.
- ✗
Deploy to a single large machine type
Why it's wrong here
A single large machine type cannot scale with traffic, so latency degrades under load and cost stays fixed during idle periods. Autoscaling with multiple replicas adjusts capacity to demand. A single large machine would suit a steady, predictable workload that consistently saturates one host.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.