easyMultiple Choice
PDE Practice Question: A company trains a custom model using TensorFlow…
A company trains a custom model using TensorFlow and wants to deploy it to Vertex AI for low-latency predictions. The model is large (2 GB). Which deployment option should they choose?
⚠ Common exam trap
Google Cloud often tests the misconception that Cloud Run or Cloud Functions can handle large models for real-time inference, ignoring their size limits, cold-start latency, and lack of native Vertex AI integration for model management and scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy to Vertex AI Endpoint with a custom container
Deploying a large (2 GB) model to Vertex AI Endpoint with a custom container allows you to package the model, its dependencies, and a serving framework (e.g., TensorFlow Serving) into a Docker image. This approach supports low-latency predictions by keeping the model loaded in memory across requests, and it can scale to handle real-time inference traffic, unlike batch or serverless options that have cold-start or size limitations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Batch Prediction job
Why it's wrong here
Batch Prediction processes asynchronous jobs against stored input and returns results to Cloud Storage, so it cannot serve low-latency online requests. It is tempting because it handles large models and bulk data cheaply, and would be correct for scheduled scoring of datasets rather than real-time inference.
- ✗
Deploy as a Cloud Function
Why it's wrong here
Cloud Functions imposes deployment package and memory limits that a 2 GB TensorFlow model exceeds, and it provides no model-serving endpoint. It is tempting for small event-driven inference, and would be correct for lightweight models triggered by events, not for low-latency predictions from a large model.
- ✓
Deploy to Vertex AI Endpoint with a custom container
Why this is correct
A 2 GB custom TensorFlow model exceeds standard pre-built container limits, so a custom container on a Vertex AI Endpoint is required. This satisfies the low-latency prediction constraint by giving full control over serving runtime and dependencies.
- ✗
Deploy to Cloud Run with minimum instances
Why it's wrong here
Cloud Run serves containerised HTTP workloads but lacks Vertex AI's model registry, versioning and prediction endpoints; a 2 GB TensorFlow model must be packaged and served manually. Cloud Run suits lightweight containerised APIs, and would fit if the requirement were a custom inference server rather than managed Vertex AI deployment.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.