PMLE Serving and Scaling Models Practice Question
A Vertex AI Endpoint hosts a model that must serve predictions with a strict 99th percentile latency under 100 ms. The model is a large TensorFlow model that processes images. During load testing, you observe that p99 latency spikes to 300 ms when batch size exceeds 1. You need to meet the latency SLO while maintaining reasonable throughput. What should you do?
⚠ Common exam trap
The trap here is assuming that more replicas or a bigger batch size will fix latency, when the spike is specifically caused by large batches.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable dynamic batching with a small maximum batch size and a short timeout.
The latency spike is caused by large batch sizes. Dynamic batching with a small maximum and short timeout limits how long requests wait and how many are combined, keeping p99 under 100 ms while still batching enough to improve throughput. Simply adding replicas or changing hardware does not resolve the batch-size-induced tail latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable dynamic batching with a small maximum batch size and a short timeout.
Why this is correct
Dynamic batching groups multiple requests into a single inference call, improving throughput, while a small maximum batch size and short timeout bound the added latency. This keeps p99 under the SLO by preventing large batches that caused the 300 ms spike, and it balances throughput and latency better than fixed batch size of 1. It directly addresses the observed trade-off.
- ✗
Increase the batch size further to amortize overhead and improve throughput.
Why it's wrong here
Increasing batch size beyond 1 already caused p99 latency to spike to 300 ms, so pushing it higher would worsen the tail latency. While larger batches can improve throughput, they directly conflict with the strict 100 ms SLO. This option misinterprets the root cause: the problem is that large batches add latency, not that overhead needs amortizing.
- ✗
Increase the number of replicas and keep batch size at 1.
Why it's wrong here
Keeping batch size at 1 avoids the latency spike but limits throughput per replica. Adding replicas increases concurrency but also cost, and it does not address the efficiency loss from processing one image at a time. This approach may meet latency but at higher cost and lower resource utilization, making it suboptimal for maintaining reasonable throughput.
- ✗
Switch to a CPU-only machine type with more cores to parallelize inference.
Why it's wrong here
CPU-only inference for a large image model is typically slower than GPU, so p99 latency would likely worsen. Adding cores can help parallelize but does not reduce the per-inference compute time enough to meet a 100 ms SLO. This change would probably increase latency and cost without solving the batch-size-induced spike.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.