hardMultiple Choice
PMLE Practice Question: A team is scaling their prototype inference model…
A team is scaling their prototype inference model to handle high-throughput requests with low latency. They use a custom container on Vertex AI Prediction. They notice that latency spikes occur under heavy load. What is the most effective strategy?
⚠ Common exam trap
PMLE often tests the reflex to throw hardware (bigger machine, GPU) at latency problems, when the expected answer is serving-layer optimization like batching and warm-up.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Optimize model serving with batching and model warm-up.
Batching groups multiple inference requests into a single forward pass, dramatically improving GPU/CPU utilization and throughput, while model warm-up pre-loads weights and initializes the runtime so the first requests after a scale-out event do not pay cold-start latency. Together they address the latency spikes under heavy load without over-provisioning. This is the standard Vertex AI Prediction optimization pattern for custom containers.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable auto-scaling with a higher minimum number of replicas.
Why it's wrong here
A higher minimum replica count only pre-warms capacity; it does not scale out during the load spike, so latency still climbs once those replicas saturate. Auto-scaling with a higher minimum suits steady baseline traffic, whereas the spike demands metric-driven scale-out on the serving replica count.
- ✓
Optimize model serving with batching and model warm-up.
Why this is correct
Batching amortises per-request overhead across concurrent inputs, while warm-up pre-loads weights and initialises the container so the first requests avoid cold-start latency. Together they directly address the latency spikes observed under heavy load on the custom container.
- ✗
Use a larger machine type with more CPUs.
Why it's wrong here
Adding CPUs cannot reduce queueing latency once the container's request-handling threads saturate; Vertex AI scales throughput by adding replicas, not by enlarging a single machine. Larger machine types suit memory-bound or GPU-bound single-replica workloads, not horizontal high-throughput serving.
- ✗
Use a GPU-based machine.
Why it's wrong here
May improve throughput but not necessarily tail latency spikes.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.