Generative AI Leader Google Cloud's Generative AI Offerings Practice Question
A startup wants to deploy a custom-tuned large language model for real-time inference on Vertex AI. They need the lowest possible latency for end users. What deployment strategy should they choose?
⚠ Common exam trap
Candidates often confuse 'lowest possible latency' with 'high throughput' or 'cost efficiency,' leading them to choose batch prediction (D) or serverless options (B) without recognizing that GPU-accelerated endpoints are specifically designed for sub-second inference.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling.
Deploying the custom-tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling is the best choice for lowest-latency real-time inference. A dedicated endpoint keeps the model loaded and ready to serve requests, avoiding the per-request startup overhead of serverless wrappers like Cloud Functions. GPU acceleration speeds up the model's forward-pass computation, and autoscaling adds or removes replicas to match traffic so that requests are not queued behind insufficient capacity. Batch prediction (D) is designed for high-throughput offline jobs and is not suitable for interactive, low-latency use cases.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Model Garden to deploy the base PaLM 2 model.
Why it's wrong here
Model Garden offers pre-built models, not custom tuned models.
- ✗
Wrap the model in a Cloud Function and invoke via HTTP.
Why it's wrong here
Cloud Functions are not optimized for model inference and have cold start issues.
- ✓
Deploy the tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling.
Why this is correct
GPU acceleration provides the compute throughput needed for low-latency token generation, while autoscaling matches capacity to demand without cold starts. Deploying to a Vertex AI endpoint keeps the model resident for real-time inference, directly satisfying the lowest-possible-latency constraint.
- ✗
Use Vertex AI Batch Prediction to process requests in batches.
Why it's wrong here
Batch prediction is not real-time; it is for offline processing.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.