easyMultiple Choice
PMLE Practice Question: A data science team has trained a TensorFlow…
A data science team has trained a TensorFlow model and wants to serve it online with minimal latency. Which Vertex AI deployment option should they use to ensure the model can handle traffic spikes without manual scaling?
⚠ Common exam trap
Google Cloud often tests the misconception that any cloud deployment with a load balancer (like Compute Engine) provides automatic scaling, but the trap here is that Vertex AI Endpoints offer managed autoscaling natively, whereas Compute Engine VMs require additional infrastructure setup and do not automatically scale without configuring managed instance groups.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the model to a Vertex AI Endpoint with automatic scaling.
Vertex AI Endpoints with automatic scaling (option B) are designed for online serving with minimal latency and can automatically adjust the number of replicas based on traffic load, handling spikes without manual intervention. This is the correct choice for a TensorFlow model requiring real-time inference and elastic scaling.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Model Garden.
Why it's wrong here
Model Garden is a catalogue of pre-trained and partner models for discovery and deployment; it does not itself provide autoscaled online serving for your custom TensorFlow artefact. It is tempting because it accelerates model selection, and would be correct when adopting an existing foundation model rather than serving your own trained one.
- ✓
Deploy the model to a Vertex AI Endpoint with automatic scaling.
Why this is correct
Deploying to a Vertex AI Endpoint with automatic scaling directly satisfies the low-latency and spike-handling constraints. Endpoints provision dedicated prediction nodes, and autoscaling adjusts replica count based on traffic, avoiding cold starts that serverless batch or custom containers would introduce.
- ✗
Use Vertex AI Batch Prediction for offline inference.
Why it's wrong here
Batch Prediction processes jobs asynchronously against stored input, returning results to a destination, so it cannot serve online requests with minimal latency. It is tempting because it suits large offline scoring runs cheaply, and would be correct for periodic bulk inference rather than real-time traffic.
- ✗
Deploy the model to a Compute Engine VM with a load balancer.
Why it's wrong here
A Compute Engine VM with a load balancer provides no autoscaling tied to model traffic; scaling remains manual, contradicting the requirement. It is tempting because VMs give full control over runtime and GPU drivers, and would be correct for custom serving stacks or workloads Vertex AI cannot host.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.