mediumMultiple Select
PMLE Practice Question: A team has trained a sentiment analysis model…
A team has trained a sentiment analysis model using PyTorch on Vertex AI Training. They now want to deploy it for online predictions with low latency. Which TWO actions should they take? (Choose 2)
⚠ Common exam trap
Google Cloud often tests the misconception that converting to TensorFlow SavedModel is required for Vertex AI, but the platform supports PyTorch natively via custom containers, making conversion an unnecessary and potentially error-prone step.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a machine type with a GPU for faster inference.
Option B is correct because deploying a PyTorch model for low-latency online predictions on Vertex AI benefits from GPU-backed machine types, which accelerate the matrix/tensor computations of deep neural network inference and reduce per-request latency. Option E is correct because Vertex AI custom containers let you serve a PyTorch model with your own web server (e.g., FastAPI) implementing the required predict/health routes, which is the standard way to deploy non-TensorFlow frameworks for online prediction. Option A is not required for low latency; model versions and A/B testing address traffic splitting and evaluation, not inference speed. Option C is wrong because batch prediction is for asynchronous, bulk scoring and does not provide the low-latency online endpoint requested. Option D is wrong because converting a PyTorch model to TensorFlow SavedModel is unnecessary and would require a framework conversion rather than serving the trained PyTorch artifact.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create multiple model versions for A/B testing.
Why it's wrong here
A/B testing splits live traffic between model versions to compare business metrics; it does nothing to reduce inference latency. Deploying the trained model to a Vertex AI Endpoint with an appropriate machine type and autoscaling addresses online prediction latency. Multiple versions suit gradual rollout or experimentation, not this requirement.
- ✓
Use a machine type with a GPU for faster inference.
Why this is correct
A GPU machine type accelerates the matrix operations underpinning PyTorch inference, directly satisfying the low-latency requirement for online predictions. Vertex AI supports attaching GPUs to prediction nodes, so tensor computations complete faster than CPU-only serving, reducing per-request response times for the sentiment model.
- ✗
Enable batch prediction instead of online prediction.
Why it's wrong here
Batch prediction writes results to storage asynchronously, so responses are not returned per request; that contradicts the low-latency online serving requirement. Batch is the right choice for scoring large datasets offline where throughput matters and immediate responses do not.
- ✗
Convert the model to TensorFlow SavedModel format.
Why it's wrong here
Vertex AI serves PyTorch models through custom containers or the PyTorch pre-built container; converting to TensorFlow SavedModel is unnecessary and discards the trained framework. Conversion would be correct when migrating to a TensorFlow serving stack that cannot host PyTorch artefacts.
- ✓
Package the model in a custom container with a web server (e.g., FastAPI).
Why this is correct
Vertex AI online prediction requires a servable artefact; packaging the PyTorch model in a custom container with a web server such as FastAPI gives the required HTTP prediction endpoint, satisfying the low-latency online serving constraint. The container must expose a health and predict route.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.