Generative AI Leader Google Cloud's Generative AI Offerings Practice Question
A company is deploying a large language model on Vertex AI for real-time inference. They observe high latency and want to optimize. They have already enabled model caching. What next step should they take to reduce latency?
⚠ Common exam trap
Google often tests the misconception that adding more GPUs or increasing batch size reduces latency for real-time inference, when in fact these optimizations primarily improve throughput and can increase per-request latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply model quantization to reduce precision
Model quantization reduces the precision of the model's weights (e.g., from FP32 to INT8), which decreases memory footprint and accelerates computation on the hardware, directly lowering inference latency. Since Vertex AI already has model caching enabled, quantization is the next logical optimization step to reduce latency without requiring additional infrastructure changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add more GPUs to the prediction endpoint
Why it's wrong here
Adding GPUs raises parallel throughput, not per-request latency; the model still executes the same sequential token generation. It is tempting because horizontal scaling suits batch or high-volume serving, but real-time latency needs a smaller distilled model, lower precision, or speculative decoding instead.
- ✗
Use a larger, more accurate model variant
Why it's wrong here
A larger model variant adds parameters and computation per token, increasing inference latency; accuracy gains do not offset the slower response. It is tempting because accuracy is often prioritised, and it would be correct where output quality outweighs speed, such as asynchronous batch processing rather than real-time serving.
- ✗
Increase the batch size for inference requests
Why it's wrong here
Increasing batch size groups more requests per inference pass, which raises throughput but adds queueing delay before each response, worsening latency for real-time serving. It is tempting because batching is a standard optimisation, and it would be correct for offline or bulk scoring where throughput matters more than per-request response time.
- ✓
Apply model quantization to reduce precision
Why this is correct
Quantization stores weights in lower precision, such as INT8, shrinking memory footprint and enabling faster matrix operations, which directly cuts inference latency. Caching is already applied, so reducing per-token compute is the remaining lever for real-time serving on Vertex AI.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.