PMLE Serving and Scaling Models Practice Question
Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?
⚠ Common exam trap
PMLE often tests whether candidates reach for infrastructure scaling (more replicas, CDN) when the correct answer is application-level caching for repeated identical or near-identical requests.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Cloud Memorystore to cache prediction results.
Cloud Memorystore (Redis) in front of the Vertex AI endpoint lets you cache prediction results keyed by user/request signature, so repeated similar requests within short windows are served from cache instead of hitting the GPU-backed model. This reduces both latency (cache hit is sub-millisecond) and cost (fewer GPU inference calls).
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to CPU-only instances to reduce cost.
Why it's wrong here
CPU-only instances remove GPU acceleration, so inference latency increases substantially on a large recommendation model. It is tempting because CPU instances cost less per hour, which suits small models or latency-tolerant batch scoring, not the low-latency GPU serving this scenario requires.
- ✗
Increase maxReplicas to handle the load without caching.
Why it's wrong here
Raising maxReplicas adds GPU instances, scaling capacity but leaving every similar request to recompute inference, so cost rises without latency reduction. It is tempting because autoscaling genuinely handles spiky concurrent load, which suits throughput-bound traffic rather than repeated identical requests.
- ✗
Set up a Cloud CDN in front of the endpoint.
Why it's wrong here
Cloud CDN caches HTTP responses at edge locations, but Vertex AI prediction requests are POST calls with user-specific payloads, which CDN does not cache. It is tempting because CDN genuinely reduces latency for cacheable static or GET content, not for dynamic model inference.
- ✓
Use Cloud Memorystore to cache prediction results.
Why this is correct
Cloud Memorystore caches repeated predictions for the same users within short windows, so identical requests bypass GPU inference entirely. This cuts both latency and GPU cost, directly addressing the repeated-request pattern while the endpoint stays GPU-backed.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.