Courseiva

Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions

A company wants to scale their generative AI application globally with low latency. Which infrastructure configuration is most suitable?

⚠ Common exam trap

Test-takers frequently confuse CDN caching with real-time inference, assuming caching can accelerate dynamic AI responses, but generative AI outputs are unique per request and cannot be pre-cached.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Multiple regional endpoints with traffic routing to the nearest region.

Deploying multiple regional endpoints with traffic routing to the nearest region minimizes latency by directing user requests to the geographically closest inference endpoint. This architecture leverages global load balancing (e.g., using Anycast DNS or HTTP(S) load balancers with backend services in multiple regions) to reduce round-trip time (RTT) and meet latency SLAs for real-time generative AI applications.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a CDN to cache responses.

    Why it's wrong here

    A CDN caches static or cacheable HTTP responses at edge locations; generative model inference is dynamic and per-prompt, so cached responses cannot serve unique requests. It is tempting because CDNs cut latency for static assets, which is the right scenario for web content delivery, not model inference.

  • ✓

    Multiple regional endpoints with traffic routing to the nearest region.

    Why this is correct

    Deploying the model to multiple regional endpoints and routing each user to the nearest region reduces network distance, satisfying the low-latency requirement for global users. A single-region endpoint would force distant users to traverse long network paths.

  • ✗

    On-premises deployment for all regions.

    Why it's wrong here

    On-premises deployment in every region cannot match the global edge footprint and managed scaling of cloud regions, so latency and elasticity suffer. It is tempting because on-premises suits strict data-residency or regulatory mandates, but that is the correct scenario, not global low-latency scaling.

  • ✗

    Single endpoint in us-central1 with high max replicas.

    Why it's wrong here

    A single regional endpoint forces distant users to traverse the globe, so latency stays high regardless of replica count. It is tempting because high max replicas handle throughput spikes, but that is the correct configuration when demand is concentrated in one region, not for global low-latency serving.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.