Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A data scientist has trained a TensorFlow model locally and wants to deploy it to Vertex AI for online predictions. The model accepts a single input tensor of shape (1, 224, 224, 3) and outputs a probability distribution over 10 classes. The data scientist wants to minimize deployment effort and ensure the model is served with low latency. What is the simplest way to deploy this model on Vertex AI?

⚠ Common exam trap

The trap here is overcomplicating the deployment by considering custom containers or format conversions when a pre-built container for TensorFlow is readily available.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Export the model as a SavedModel and upload it to Vertex AI Model Registry, then deploy to an endpoint with a pre-built TensorFlow Serving container.

The simplest and most efficient way to deploy a standard TensorFlow model on Vertex AI for online predictions is to export it as a SavedModel and use the pre-built TensorFlow Serving container. This requires no custom serving code, leverages Vertex AI's managed infrastructure, and provides low-latency predictions. It also integrates with Vertex AI Model Registry for versioning and monitoring.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the model to ONNX format and use a custom container with ONNX Runtime for serving.

    Why it's wrong here

    While ONNX Runtime can serve models efficiently, converting to ONNX and building a custom container adds unnecessary complexity. Vertex AI has native support for TensorFlow SavedModel, so using ONNX here is not the simplest approach and increases deployment effort.

  • ✗

    Use Vertex AI Batch Prediction with a pre-built TensorFlow container, and then set up a Cloud Function to serve online requests.

    Why it's wrong here

    Batch Prediction is designed for offline, large-scale predictions, not for low-latency online serving. Using a Cloud Function to call batch prediction would be inefficient and not suitable for real-time requests. This approach does not meet the requirement for online predictions with low latency.

  • ✓

    Export the model as a SavedModel and upload it to Vertex AI Model Registry, then deploy to an endpoint with a pre-built TensorFlow Serving container.

    Why this is correct

    Vertex AI supports deploying TensorFlow SavedModels directly using a pre-built TensorFlow Serving container. This requires no custom code, and the serving container handles the model signature. It is the simplest and most efficient way to deploy a standard TensorFlow model for online predictions with low latency.

  • ✗

    Package the model in a Docker container with a Flask app that loads the model and exposes a REST API, then deploy as a custom container on Vertex AI.

    Why it's wrong here

    Building a custom Flask container requires writing and maintaining serving code, which is more effort than using a pre-built container. While it offers flexibility, it is not the simplest method for a standard TensorFlow model. It also may introduce latency if not optimized.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.