Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

A startup wants to deploy a custom-tuned large language model for real-time inference on Vertex AI. They need the lowest possible latency for end users. What deployment strategy should they choose?

⚠ Common exam trap

Candidates often confuse 'lowest possible latency' with 'high throughput' or 'cost efficiency,' leading them to choose batch prediction (D) or serverless options (B) without recognizing that GPU-accelerated endpoints are specifically designed for sub-second inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy the tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling.

Deploying the custom-tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling is the best choice for lowest-latency real-time inference. A dedicated endpoint keeps the model loaded and ready to serve requests, avoiding the per-request startup overhead of serverless wrappers like Cloud Functions. GPU acceleration speeds up the model's forward-pass computation, and autoscaling adds or removes replicas to match traffic so that requests are not queued behind insufficient capacity. Batch prediction (D) is designed for high-throughput offline jobs and is not suitable for interactive, low-latency use cases.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Vertex AI Model Garden to deploy the base PaLM 2 model.

    Why it's wrong here

    Model Garden offers pre-built models, not custom tuned models.

  • ✗

    Wrap the model in a Cloud Function and invoke via HTTP.

    Why it's wrong here

    Cloud Functions are not optimized for model inference and have cold start issues.

  • ✓

    Deploy the tuned model to a Vertex AI endpoint with GPU acceleration and autoscaling.

    Why this is correct

    GPU acceleration provides the compute throughput needed for low-latency token generation, while autoscaling matches capacity to demand without cold starts. Deploying to a Vertex AI endpoint keeps the model resident for real-time inference, directly satisfying the lowest-possible-latency constraint.

  • ✗

    Use Vertex AI Batch Prediction to process requests in batches.

    Why it's wrong here

    Batch prediction is not real-time; it is for offline processing.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.