Courseiva

PMLE Serving and Scaling Models Practice Question

Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?

⚠ Common exam trap

The trap is selecting Cloud CDN because it is a caching service, but CDN caches HTTP responses at the edge and is not suitable for caching dynamic API prediction results that require custom key logic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Memorystore for Redis

Cloud Memorystore for Redis is an in-memory data store that provides sub-millisecond latency, making it ideal for caching prediction responses. By caching identical prediction requests, you can reduce the number of calls to the Vertex AI endpoint, lowering latency and cost. Redis supports key-value storage with TTL, perfect for caching.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Cloud Memorystore for Redis

    Why this is correct

    Memorystore for Redis provides a low-latency in-memory store that sits in front of the endpoint, letting identical feature vectors return cached predictions instead of re-invoking the model. This directly cuts both response latency and per-prediction cost, matching the stem's caching goal.

  • ✗

    Cloud CDN

    Why it's wrong here

    Cloud CDN caches HTTP responses at edge locations keyed by URL, so it cannot match identical prediction payloads sent to a Vertex AI endpoint. It is tempting because it reduces latency, but it fronts web content, not gRPC or REST inference requests.

  • ✗

    Bigtable

    Why it's wrong here

    Bigtable stores time-series and analytical data at petabyte scale, but offers no response-caching primitive for identical prediction inputs. The tempting fit is its low-latency key lookups, yet the scenario needs a managed cache layer keyed on request payloads, which Memorystore or a dedicated caching tier provides.

  • ✗

    Cloud Storage with object versioning

    Why it's wrong here

    Cloud Storage with object versioning retains multiple generations of an object; it provides no low-latency key-value lookup for identical prediction inputs. It is tempting as durable storage, but caching repeated inference responses requires Memorystore for Redis or a similar in-memory cache.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.