Generative AI Leader Fundamentals of Generative AI Practice Question
Which THREE of the following are key considerations when deploying a generative AI model in a production environment with strict latency requirements? (Choose three.)
⚠ Common exam trap
Google Cloud often tests the distinction between latency and throughput, so the trap here is that candidates confuse batch size (which improves throughput) with latency reduction, or assume larger models always yield better performance without considering inference speed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement speculative decoding to generate candidate tokens with a smaller draft model and verify with the large model.
Option B is correct because speculative decoding uses a smaller, faster draft model to propose multiple candidate tokens that the larger target model then verifies in parallel, reducing the number of expensive forward passes and lowering end-to-end latency while preserving output quality. Option C is correct because quantizing weights and activations to lower precision such as int8 shrinks memory bandwidth demands and enables faster matrix multiplications on supported hardware, directly cutting per-token inference time. Option D is correct because caching the key-value tensors from prior decoding steps avoids recomputing attention over the entire sequence for each new token, which is essential for keeping autoregressive generation latency low. Option A is not appropriate because the largest model variant increases compute and memory cost, worsening latency rather than meeting strict requirements. Option E is not appropriate because increasing batch size improves throughput and GPU utilization but typically raises per-request latency, which conflicts with strict latency goals.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the largest model variant available to ensure highest quality.
Why it's wrong here
Larger variants add parameters and decoding time per token, pushing latency beyond strict real-time budgets. It is tempting because bigger models usually raise output quality, and would be the right pick when accuracy matters and latency targets are loose.
- ✓
Implement speculative decoding to generate candidate tokens with a smaller draft model and verify with the large model.
Why this is correct
Speculative decoding attacks latency directly: a smaller draft model proposes candidate tokens cheaply, and the large model verifies them in parallel, preserving output quality while cutting sequential decoding steps. This meets the strict latency requirement without retraining or accuracy loss.
- ✓
Use model quantization (e.g., int8) to reduce precision and speed up matrix multiplications.
Why this is correct
Quantisation stores weights and activations at lower precision, such as int8, so matrix multiplications execute faster and consume less memory bandwidth. This directly reduces per-token inference latency, satisfying the strict latency constraint in the stem.
- ✓
Cache the key-value caches from previous decoding steps to avoid redundant computation.
Why this is correct
Autoregressive decoding recomputes attention over all prior tokens each step. Caching keys and values from previous steps eliminates that redundant work, so each new token requires only incremental computation, directly lowering latency under the strict requirement.
- ✗
Increase the inference batch size to maximize GPU utilization.
Why it's wrong here
Batching requests together delays each individual response until the batch fills, raising per-request latency under strict real-time requirements. It is tempting because larger batches raise GPU throughput and cut cost per token, which suits offline or throughput-bound workloads rather than latency-sensitive ones.
Go deeper
Related to this question
About these practice questions
This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.