easyMultiple Choice
Scaling a TensorFlow Model to Production
A team has a trained TensorFlow model running locally and wants to deploy it for low-latency online predictions on Google Cloud. Which service should they use?
⚠ Common exam trap
Google Cloud often tests the distinction between training and prediction services, and the trap here is that candidates may confuse AI Platform Training (which is for model training) with AI Platform Prediction (now part of Vertex AI), or assume that any serverless compute like Cloud Run or Cloud Functions can handle ML inference without considering the need for GPU/TPU support and optimized serving infrastructure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Vertex AI Prediction
Vertex AI Prediction is the correct choice because it is a fully managed service designed specifically for deploying trained ML models for online (real-time) prediction with low latency. It supports importing TensorFlow SavedModel artifacts and automatically scales the serving infrastructure, including GPU/TPU support, to handle request traffic while providing built-in monitoring and explainability features.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Vertex AI Prediction
Why this is correct
Vertex AI Prediction provides a managed endpoint with low-latency online serving, letting the trained TensorFlow model be deployed as a container and queried in real time. It satisfies the low-latency online prediction constraint that batch or local serving cannot meet.
- ✗
AI Platform Training
Why it's wrong here
AI Platform Training runs training and hyperparameter tuning jobs; it does not host a persistent prediction endpoint for low-latency serving. It tempts because it is a Vertex AI component that accepts TensorFlow models, but its purpose is producing trained artefacts, not serving online requests. Prediction requires a deployed model endpoint.
- ✗
Cloud Run
Why it's wrong here
Cloud Run serves containerised HTTP applications but provides no model registry, versioning, or prediction endpoint; the team would have to build serving infrastructure themselves. It tempts because it offers low-latency autoscaling HTTP. Vertex AI Prediction is purpose-built for deploying trained models with managed online prediction endpoints.
- ✗
Cloud Functions
Why it's wrong here
Cloud Functions runs event-driven single-purpose code with cold starts and execution time limits, unsuited to sustained low-latency model inference. It tempts because it is serverless and simple to invoke over HTTP. Vertex AI Prediction provides managed model deployment, versioning, and dedicated prediction endpoints instead.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.