easyMultiple Choice
PMLE Practice Question: Serve a model with strict latency requirements…
A company needs to serve a model with strict latency requirements (<100ms). They are using Vertex AI Prediction with CPU. During testing, latency is 150ms. What should they do?
⚠ Common exam trap
Test-takers frequently confuse throughput optimization (batching or scaling replicas) with latency reduction, failing to recognize that GPUs directly address compute-bound latency while CPU-based solutions cannot meet strict sub-100ms requirements for complex models.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch to a GPU machine type
The model's latency of 150ms exceeds the 100ms requirement. Switching to a GPU machine type (Option D) is correct because GPUs are optimized for parallel computation, significantly reducing inference latency for many ML models, especially deep learning models, compared to CPUs. Vertex AI Prediction supports GPU machine types, and this change directly addresses the latency bottleneck without altering the model or its serving configuration.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable batching to improve throughput
Why it's wrong here
Batching groups requests to raise throughput, but it adds queueing delay before inference starts, pushing per-request latency further above 100ms. It tempts because batching is the standard lever for cost and throughput optimisation, yet the stem's constraint is latency, not throughput.
- ✗
Use a smaller machine type with more replicas
Why it's wrong here
More replicas of a smaller machine type add horizontal capacity, not per-request speed; each individual inference still runs on slower hardware. It tempts because scaling out is the usual remedy for overloaded CPU serving, but the stem's 150ms figure reflects single-request compute time, not queue contention.
- ✗
Export the model to TensorFlow Lite
Why it's wrong here
TensorFlow Lite targets mobile and edge deployment, and converting a model from another framework risks unsupported operators and accuracy loss. It tempts because lightweight runtimes do cut inference latency, but the stem's CPU serving latency is addressed by GPU acceleration, not by changing the runtime format.
- ✓
Switch to a GPU machine type
Why this is correct
GPUs accelerate the matrix multiplications dominating neural network inference, cutting compute time well below the 100ms threshold that CPU inference currently exceeds at 150ms. For latency-bound serving, GPU machine types provide the parallel throughput needed, satisfying the stem's strict sub-100ms constraint where CPU cannot.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.