Courseiva

Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions

A global news agency is using a generative AI model to summarize breaking news articles in real-time. The model is deployed on Vertex AI across multiple regions (us-central1, europe-west4, asia-southeast1) for low latency worldwide. The agency has a Service Level Objective (SLO) of 99.9% availability and p99 latency under 2 seconds. Recently, during a major event, traffic spiked 10x, and the europe-west4 region experienced latency spikes over 5 seconds and some 503 errors. The team suspects the regional endpoint is under-provisioned. Which combination of actions should they take to meet the SLO consistently?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable the global endpoint feature in Vertex AI with automatic traffic splitting, and increase the minimum replicas for each regional endpoint

It enables the global endpoint feature with automatic traffic splitting, allowing traffic to be routed to healthy regions and providing failover. Additionally, increasing minimum replicas per region ensures each regional endpoint has baseline capacity to handle spikes, preventing under-provisioning. Option B only increases max replicas in europe-west4, which does not address traffic shifts, and reducing min replicas elsewhere risks capacity issues. Option C suggests Cloud CDN, which is for static content, not model inference. Option D configures a global load balancer with a single endpoint, which does not optimally use Vertex AI's regional endpoints and may not meet latency SLO.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable the global endpoint feature in Vertex AI with automatic traffic splitting, and increase the minimum replicas for each regional endpoint

    Why this is correct

    Global endpoint distributes traffic and increases capacity; higher min replicas prevent cold starts during spikes.

  • ✗

    Increase the maximum replicas for the europe-west4 endpoint and reduce the min replicas in other regions

    Why it's wrong here

    Reducing min replicas in other regions could cause latency issues there during spikes.

  • ✗

    Implement Cloud CDN caching for common summaries and reduce the number of regions to two

    Why it's wrong here

    Cloud CDN is not suitable for dynamic model inference; reducing regions may increase latency for some users.

  • ✗

    Configure a global load balancer with a single Vertex AI endpoint and increase max replicas globally

    Why it's wrong here

    A single endpoint may not reduce latency; still need sufficient replicas in each region.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.