Courseiva
easyMultiple Choice

PDE Practice Question: Deploy a trained model for real-time predictions…

A company needs to deploy a trained model for real-time predictions with low latency. Which Vertex AI resource should they use?

⚠ Common exam trap

Google Cloud often tests the distinction between batch and online prediction, and the trap here is that candidates confuse Vertex AI Batch Prediction (which is for offline, large-scale inference) with the real-time serving capability of Vertex AI Endpoints, leading them to select option B.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Vertex AI Endpoints

Vertex AI Endpoints are designed for online prediction, providing a managed service that hosts models for real-time inference with low latency. They automatically scale resources and handle traffic routing, making them the correct choice for deploying a trained model that needs to respond to individual prediction requests quickly.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud TPU

    Why it's wrong here

    Cloud TPUs are accelerator hardware for training and running large models, not a managed serving endpoint, so they provide no request-handling or autoscaling interface for real-time predictions. They are tempting because they reduce inference compute time, and would be the right choice for accelerating training or bulk inference workloads.

  • ✗

    Vertex AI Batch Prediction

    Why it's wrong here

    Batch Prediction processes large jobs asynchronously from stored input, returning results later rather than serving per-request responses, so it cannot meet low-latency real-time serving. It is tempting because it is the correct Vertex AI resource for high-volume offline scoring where latency is irrelevant.

  • ✓

    Vertex AI Endpoints

    Why this is correct

    Endpoints host trained models for online serving, exposing a REST/gRPC interface that returns synchronous predictions per request. This satisfies the low-latency real-time constraint, unlike batch prediction, which processes asynchronous jobs over stored data and cannot serve individual requests on demand.

  • ✗

    Cloud Run

    Why it's wrong here

    Cloud Run hosts containerised HTTP services but provides no model registry, versioning or prediction endpoint, so it does not deliver a managed Vertex AI deployment. It is tempting because it scales containers automatically, and would be correct for serving a custom containerised application rather than a registered model.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.