Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions
A global news agency is using a generative AI model to summarize breaking news articles in real-time. The model is deployed on Vertex AI across multiple regions (us-central1, europe-west4, asia-southeast1) for low latency worldwide. The agency has a Service Level Objective (SLO) of 99.9% availability and p99 latency under 2 seconds. Recently, during a major event, traffic spiked 10x, and the europe-west4 region experienced latency spikes over 5 seconds and some 503 errors. The team suspects the regional endpoint is under-provisioned. Which combination of actions should they take to meet the SLO consistently?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable the global endpoint feature in Vertex AI with automatic traffic splitting, and increase the minimum replicas for each regional endpoint
It enables the global endpoint feature with automatic traffic splitting, allowing traffic to be routed to healthy regions and providing failover. Additionally, increasing minimum replicas per region ensures each regional endpoint has baseline capacity to handle spikes, preventing under-provisioning. Option B only increases max replicas in europe-west4, which does not address traffic shifts, and reducing min replicas elsewhere risks capacity issues. Option C suggests Cloud CDN, which is for static content, not model inference. Option D configures a global load balancer with a single endpoint, which does not optimally use Vertex AI's regional endpoints and may not meet latency SLO.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable the global endpoint feature in Vertex AI with automatic traffic splitting, and increase the minimum replicas for each regional endpoint
Why this is correct
Global endpoint distributes traffic and increases capacity; higher min replicas prevent cold starts during spikes.
- ✗
Increase the maximum replicas for the europe-west4 endpoint and reduce the min replicas in other regions
Why it's wrong here
Reducing min replicas in other regions could cause latency issues there during spikes.
- ✗
Implement Cloud CDN caching for common summaries and reduce the number of regions to two
Why it's wrong here
Cloud CDN is not suitable for dynamic model inference; reducing regions may increase latency for some users.
- ✗
Configure a global load balancer with a single Vertex AI endpoint and increase max replicas globally
Why it's wrong here
A single endpoint may not reduce latency; still need sufficient replicas in each region.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.