Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions
Exhibit
Refer to the exhibit.
```json
{
"deployment": {
"machineType": "n1-highmem-16",
"minReplicaCount": 1,
"maxReplicaCount": 5,
"accelerator": {
"acceleratorType": "NVIDIA_TESLA_T4",
"acceleratorCount": 1
},
"trafficSplit": {"default": 100}
}
}
```A company deployed a large language model on Vertex AI using the configuration shown in the exhibit. During peak usage, users report high latency. Which change is most likely to improve latency?
⚠ Common exam trap
The Generative AI Leader exam often tests the misconception that upgrading hardware (GPU memory or type) is the primary fix for latency, when in fact scaling out replicas is the more direct solution for handling concurrent request load.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase minReplicaCount to 3.
Increasing minReplicaCount to 3 ensures that at least three instances of the model are always running and ready to serve requests. This reduces cold-start latency and distributes the load across multiple replicas, directly addressing high latency during peak usage by providing more concurrent serving capacity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove the accelerator to simplify deployment.
Why it's wrong here
Removing the accelerator forces inference onto CPU, which dramatically slows token generation and worsens latency. Simplifying deployment is the real aim of that change, and it would suit a low-traffic development or testing environment where latency is not a concern.
- ✓
Increase minReplicaCount to 3.
Why this is correct
Vertex AI scales replicas between minReplicaCount and maxReplicaCount; raising the minimum to 3 keeps additional instances warm, so peak traffic is spread across more replicas and per-request latency falls. Cold-start delays from scaling up are avoided.
- ✗
Switch to a GPU with more memory, such as NVIDIA_TESLA_A100.
Why it's wrong here
Swapping to an A100 with more memory addresses model capacity, not the throughput bottleneck causing peak-usage latency; memory headroom does not by itself raise request-processing speed. It would be the right change when the model is too large to fit on the existing accelerator.
- ✗
Change machineType to n1-standard-4 to reduce cost.
Why it's wrong here
A smaller n1-standard-4 machine type reduces compute capacity, which increases inference latency under peak load rather than improving it. Cost reduction is the actual purpose of that change, making it appropriate when the priority is lowering spend on a lightly used, latency-tolerant deployment.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.