PMLE Serving and Scaling Models Practice Question
Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?
⚠ Common exam trap
Many exam-takers confuse perceived latency improvements from streaming with actual compute reductions, when the real gains come from batching and KV-cache memory management.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement continuous batching and paged attention in the serving container.
For autoregressive LLM serving, the key bottlenecks are GPU underutilization during decoding and memory fragmentation in the key-value cache. Continuous batching keeps the GPU busy by admitting new requests into the current batch, while paged attention stores the KV cache in fixed-size blocks to avoid fragmentation and support more concurrent sequences. These techniques reduce per-token latency and improve throughput, making them the right optimization for long prompts on a cost-conscious deployment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a smaller quantized version of the model and serve it with a batch size of one.
Why it's wrong here
Quantization reduces memory footprint and can speed up matrix operations, but serving with a batch size of one wastes GPU parallelism and does not exploit continuous batching. A smaller model may also reduce quality. This combination does not specifically target the long-prompt latency, and batch size one can increase cost per token because the GPU is underutilized during decoding.
- ✓
Implement continuous batching and paged attention in the serving container.
Why this is correct
Continuous batching allows the server to add new requests to an in-flight batch at each decoding step, improving GPU utilization and throughput. Paged attention manages the key-value cache in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together they lower per-token latency and increase the number of requests served per GPU, which directly addresses long-prompt inference cost and latency.
- ✗
Increase maxReplicaCount and rely on autoscaling to handle long prompts.
Why it's wrong here
Adding replicas increases aggregate capacity but does not reduce the latency of a single request. Each replica still processes long prompts with the same per-token cost, and autoscaling reacts after load increases. Without efficient batching and memory management, more replicas simply multiply cost without improving the time to first token or per-token latency for a given request.
- ✗
Enable response streaming and return tokens as they are generated.
Why it's wrong here
Streaming improves perceived latency by delivering tokens incrementally, but it does not reduce the total compute or the time to generate the full response. For long prompts, the prefill phase still processes all input tokens, and the decoding phase still runs token by token. Streaming changes how results are delivered, not how fast the model computes, so it does not address the underlying latency for long prompts.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.