PMLE Serving and Scaling Models Practice Question
Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?
⚠ Common exam trap
The trap is selecting Cloud CDN because it is a caching service, but CDN caches HTTP responses at the edge and is not suitable for caching dynamic API prediction results that require custom key logic.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Memorystore for Redis
Cloud Memorystore for Redis is an in-memory data store that provides sub-millisecond latency, making it ideal for caching prediction responses. By caching identical prediction requests, you can reduce the number of calls to the Vertex AI endpoint, lowering latency and cost. Redis supports key-value storage with TTL, perfect for caching.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Cloud Memorystore for Redis
Why this is correct
Memorystore for Redis provides a low-latency in-memory store that sits in front of the endpoint, letting identical feature vectors return cached predictions instead of re-invoking the model. This directly cuts both response latency and per-prediction cost, matching the stem's caching goal.
- ✗
Cloud CDN
Why it's wrong here
Cloud CDN caches HTTP responses at edge locations keyed by URL, so it cannot match identical prediction payloads sent to a Vertex AI endpoint. It is tempting because it reduces latency, but it fronts web content, not gRPC or REST inference requests.
- ✗
Bigtable
Why it's wrong here
Bigtable stores time-series and analytical data at petabyte scale, but offers no response-caching primitive for identical prediction inputs. The tempting fit is its low-latency key lookups, yet the scenario needs a managed cache layer keyed on request payloads, which Memorystore or a dedicated caching tier provides.
- ✗
Cloud Storage with object versioning
Why it's wrong here
Cloud Storage with object versioning retains multiple generations of an object; it provides no low-latency key-value lookup for identical prediction inputs. It is tempting as durable storage, but caching repeated inference responses requires Memorystore for Redis or a similar in-memory cache.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.