PMLE Serving and Scaling Models Practice Question
Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?
⚠ Common exam trap
A common mix-up: candidates confuse caching at the edge (CDN) with caching at the application layer (Memorystore), assuming any cache service works for dynamic API responses, but Cloud CDN cannot cache POST requests or application-specific payloads without significant configuration and still lacks the fine-grained key-value semantics needed for identical prediction requests.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Memorystore
Cloud Memorystore (B) is correct because it provides a managed in-memory cache (Redis or Memcached) that can store the results of identical prediction requests, reducing latency and load on the prediction endpoint. By caching responses keyed on the user profile or request parameters, subsequent identical requests can be served directly from Memorystore in microseconds, avoiding redundant model inference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud CDN
Why it's wrong here
Cloud CDN caches HTTP responses at edge points of presence, keyed by URL; it cannot cache arbitrary model predictions keyed by request payload unless the endpoint returns cacheable HTTP responses. It is correct for accelerating static or cacheable web content delivery.
- ✓
Cloud Memorystore
Why this is correct
Cloud Memorystore provides a managed Redis or Memcached layer that stores prediction results keyed by request parameters, so identical requests return cached responses without invoking the endpoint. This cuts both latency and endpoint load, matching the repeated-prediction pattern described.
- ✗
Cloud Spanner
Why it's wrong here
Cloud Spanner is a horizontally scalable relational database, so it stores and queries rows rather than serving cached prediction responses at the edge. It would be the right pick when the requirement is strongly consistent, globally distributed transactional data, not repeated identical inference results.
- ✗
BigQuery
Why it's wrong here
BigQuery is an analytics warehouse for running SQL over large datasets, not a low-latency response cache in front of a prediction endpoint. It would be the right choice when the requirement is storing and analysing historical prediction or event data at scale.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.