A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?
Cloud Memorystore offers a managed in-memory cache that the serving path can query before invoking the model, returning stored responses for repeated identical requests. This reduces endpoint invocations, lowering both latency and cost as the stem requires.
Why this answer
Cloud Memorystore provides a managed Redis or Memcached in-memory cache that can store frequent identical prediction requests and their responses, dramatically reducing latency and backend load. It is the standard Google Cloud service for application-level caching in front of Vertex AI endpoints.
Exam trap
PMLE often tests the distinction between caching layers — candidates may pick Cloud CDN for 'caching' without realizing CDN only caches HTTP responses at the edge and cannot key on arbitrary prediction payloads.
How to eliminate wrong answers
Option A is wrong because Cloud Bigtable is a wide-column NoSQL database optimized for high-throughput analytics and time-series data, not for low-latency key-value caching of prediction responses. Option B is wrong because Cloud CDN caches HTTP content at edge locations for static assets, not dynamic prediction API responses keyed by request payload. Option D is wrong because Cloud SQL is a relational database with millisecond-to-tens-of-milliseconds latency, far slower than in-memory caching and not designed for high-QPS cache workloads.