easyMultiple Choice
PDE Practice Question: Deploy a trained model for real-time predictions…
A company needs to deploy a trained model for real-time predictions with low latency. Which Vertex AI resource should they use?
⚠ Common exam trap
Google Cloud often tests the distinction between batch and online prediction, and the trap here is that candidates confuse Vertex AI Batch Prediction (which is for offline, large-scale inference) with the real-time serving capability of Vertex AI Endpoints, leading them to select option B.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Vertex AI Endpoints
Vertex AI Endpoints are designed for online prediction, providing a managed service that hosts models for real-time inference with low latency. They automatically scale resources and handle traffic routing, making them the correct choice for deploying a trained model that needs to respond to individual prediction requests quickly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud TPU
Why it's wrong here
Cloud TPUs are accelerator hardware for training and running large models, not a managed serving endpoint, so they provide no request-handling or autoscaling interface for real-time predictions. They are tempting because they reduce inference compute time, and would be the right choice for accelerating training or bulk inference workloads.
- ✗
Vertex AI Batch Prediction
Why it's wrong here
Batch Prediction processes large jobs asynchronously from stored input, returning results later rather than serving per-request responses, so it cannot meet low-latency real-time serving. It is tempting because it is the correct Vertex AI resource for high-volume offline scoring where latency is irrelevant.
- ✓
Vertex AI Endpoints
Why this is correct
Endpoints host trained models for online serving, exposing a REST/gRPC interface that returns synchronous predictions per request. This satisfies the low-latency real-time constraint, unlike batch prediction, which processes asynchronous jobs over stored data and cannot serve individual requests on demand.
- ✗
Cloud Run
Why it's wrong here
Cloud Run hosts containerised HTTP services but provides no model registry, versioning or prediction endpoint, so it does not deliver a managed Vertex AI deployment. It is tempting because it scales containers automatically, and would be correct for serving a custom containerised application rather than a registered model.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.