Courseiva

PMLE Serving and Scaling Models Practice Question

Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?

⚠ Common exam trap

Many exam-takers confuse perceived latency improvements from streaming with actual compute reductions, when the real gains come from batching and KV-cache memory management.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement continuous batching and paged attention in the serving container.

For autoregressive LLM serving, the key bottlenecks are GPU underutilization during decoding and memory fragmentation in the key-value cache. Continuous batching keeps the GPU busy by admitting new requests into the current batch, while paged attention stores the KV cache in fixed-size blocks to avoid fragmentation and support more concurrent sequences. These techniques reduce per-token latency and improve throughput, making them the right optimization for long prompts on a cost-conscious deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a smaller quantized version of the model and serve it with a batch size of one.

    Why it's wrong here

    Quantization reduces memory footprint and can speed up matrix operations, but serving with a batch size of one wastes GPU parallelism and does not exploit continuous batching. A smaller model may also reduce quality. This combination does not specifically target the long-prompt latency, and batch size one can increase cost per token because the GPU is underutilized during decoding.

  • ✓

    Implement continuous batching and paged attention in the serving container.

    Why this is correct

    Continuous batching allows the server to add new requests to an in-flight batch at each decoding step, improving GPU utilization and throughput. Paged attention manages the key-value cache in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together they lower per-token latency and increase the number of requests served per GPU, which directly addresses long-prompt inference cost and latency.

  • ✗

    Increase maxReplicaCount and rely on autoscaling to handle long prompts.

    Why it's wrong here

    Adding replicas increases aggregate capacity but does not reduce the latency of a single request. Each replica still processes long prompts with the same per-token cost, and autoscaling reacts after load increases. Without efficient batching and memory management, more replicas simply multiply cost without improving the time to first token or per-token latency for a given request.

  • ✗

    Enable response streaming and return tokens as they are generated.

    Why it's wrong here

    Streaming improves perceived latency by delivering tokens incrementally, but it does not reduce the total compute or the time to generate the full response. For long prompts, the prefill phase still processes all input tokens, and the decoding phase still runs token by token. Streaming changes how results are delivered, not how fast the model computes, so it does not address the underlying latency for long prompts.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.