Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

Exhibit

Refer to the exhibit.
```
$ gcloud ai endpoints list --region=us-central1
ENDPOINT_ID: 123456789
DISPLAY_NAME: my-endpoint
MODEL: projects/my-project/locations/us-central1/models/987654321
DEPLOYED_MODELS: projects/my-project/locations/us-central1/models/987654321@1
MACHINE_TYPE: n1-standard-2
MIN_REPLICA_COUNT: 1
MAX_REPLICA_COUNT: 5
```

Refer to the exhibit. A team has deployed a model to an endpoint with the configuration shown. They notice that during peak traffic, the endpoint frequently returns 429 (Too Many Requests) errors. Which action should they take to resolve this issue?

⚠ Common exam trap

Watch out — candidates often confuse vertical scaling (bigger machine type) with horizontal scaling (more replicas); 429 errors are a concurrency/throughput signal, not a memory or CPU-per-instance signal.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase MIN_REPLICA_COUNT to 5

The 429 (Too Many Requests) errors indicate the endpoint is receiving more concurrent requests than its current replica capacity can handle. Increasing MIN_REPLICA_COUNT raises the baseline number of serving replicas, so the endpoint has more capacity to absorb peak traffic without throttling. This directly addresses the throughput bottleneck rather than changing the instance shape or disabling scaling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Change MACHINE_TYPE to n1-highmem-4

    Why it's wrong here

    429 responses signal request-rate throttling, not memory exhaustion, so changing the machine type leaves the quota constraint untouched; raising the replica count or quota addresses it. It is tempting because larger machines help memory-bound models, and n1-highmem-4 would be correct for a model exceeding its current instance's RAM.

  • ✓

    Increase MIN_REPLICA_COUNT to 5

    Why this is correct

    429 errors indicate the endpoint's replicas cannot absorb peak request volume. Raising MIN_REPLICA_COUNT to 5 keeps more serving replicas warm, increasing concurrent request capacity and distributing traffic so the quota per replica is no longer exceeded.

  • ✗

    Decrease MAX_REPLICA_COUNT to 1

    Why it's wrong here

    Lowering MAX_REPLICA_COUNT to one removes the endpoint's ability to spread load across replicas, worsening throttling under peak traffic; more replicas are needed. It is tempting because reducing replicas appears to cut contention, and it would be correct for trimming cost during predictable low-traffic periods.

  • ✗

    Disable autoscaling by setting MIN_REPLICA_COUNT equals MAX_REPLICA_COUNT

    Why it's wrong here

    Pinning MIN_REPLICA_COUNT to MAX_REPLICA_COUNT disables autoscaling, freezing capacity so peak traffic still exceeds the fixed replica ceiling and returns 429s. It is tempting because fixed replica counts give predictable latency and cost, and it would be correct for steady, well-characterised workloads needing consistent performance.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.