Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions
A large enterprise runs a generative AI solution serving millions of daily inference requests. To reduce costs, they propose using serverless endpoints (Vertex AI Prediction) with a custom container, but they notice high latency during cold starts. Which strategy best addresses this problem while minimizing cost?
⚠ Common exam trap
Google Cloud often tests the misconception that prewarming via idle timeout is a configurable parameter in serverless ML services, but in Vertex AI Prediction, the idle timeout is fixed and not user-adjustable, making minimum replicas the correct approach.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set a minimum number of replicas to maintain a baseline of always-on instances.
Setting a minimum number of replicas ensures that a baseline of always-on instances is maintained, eliminating cold starts for the majority of requests. This directly addresses the latency spike caused by container initialization and model loading in serverless endpoints, while the cost impact is limited to the minimum replicas rather than scaling all instances.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Set a minimum number of replicas to maintain a baseline of always-on instances.
Why this is correct
Correct. Setting a minimum number of replicas ensures a baseline of always-on instances, which eliminates cold starts for the majority of requests. This directly addresses the latency spike caused by container initialization and model loading, and the cost is limited to the minimum replicas rather than scaling all instances.
- ✗
Upgrade to GPU-accelerated machines for all replicas.
Why it's wrong here
Incorrect. Upgrading to GPU-accelerated machines increases cost significantly and does not directly solve the cold start issue. While GPUs may reduce inference latency per request, they do not prevent the initial delay when an instance is first created.
- ✗
Implement client-side request batching to reduce the number of inference calls.
Why it's wrong here
Incorrect. Client-side request batching reduces the number of inference calls but does not address cold start latency. Batching may help with throughput, but individual requests still face cold start delays if no instances are ready.
- ✗
Use prewarmed containers by setting an idle timeout to keep instances alive.
Why it's wrong here
Incorrect. Prewarmed containers via idle timeout is not configurable in Vertex AI Prediction; the idle timeout is fixed. Additionally, this approach would keep instances alive only briefly and not guarantee availability for infrequent traffic patterns, increasing cost without fully eliminating cold starts.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.