Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

An application serving predictions from a Vertex AI endpoint receives many identical requests within a short time window. The team notices redundant computation and wants to cache responses to reduce latency and cost. What is the recommended solution?

⚠ Common exam trap

A common mix-up: candidates assume Vertex AI has a native caching feature (like an `enable_cache` flag) because other Google Cloud services offer caching, but Vertex AI endpoints require an external cache layer like Cloud Memorystore for Redis.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement a cache layer using Cloud Memorystore for Redis, hashing prediction requests.

Vertex AI does not provide built-in request caching; instead, the recommended pattern is to implement an external cache like Cloud Memorystore for Redis. By hashing the prediction request payload and using it as a cache key, identical requests within the short time window can be served from Redis, eliminating redundant model inference and reducing both latency and cost.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model on a larger machine type to handle duplicate requests faster.

    Why it's wrong here

    Scaling to a larger machine type increases throughput for concurrent requests but does not eliminate redundant computation; the same predictions will still be recomputed for each identical request within the window. This option is tempting because vertical scaling is a common remedy for latency under high load, and it would be correct if the bottleneck were insufficient compute capacity for unique requests rather than repeated processing of duplicates.

  • ✗

    Enable Vertex AI endpoint caching by setting the `enable_cache` flag.

    Why it's wrong here

    Vertex AI endpoints expose no `enable_cache` flag; prediction caching is not a configurable endpoint setting, so this option cannot be implemented. It is tempting because it sounds like a native one-line fix, and would be right if the platform actually offered a documented caching parameter for online prediction.

  • ✓

    Implement a cache layer using Cloud Memorystore for Redis, hashing prediction requests.

    Why this is correct

    Cloud Memorystore for Redis supplies a low-latency shared cache; hashing each prediction request produces a deterministic key so identical requests return the stored response instead of recomputing. This directly removes the redundant computation and cost the stem describes.

  • ✗

    Use Cloud CDN in front of the endpoint.

    Why it's wrong here

    Cloud CDN caches static HTTP content at edge locations and does not cache Vertex AI prediction responses, so identical requests still reach the endpoint and trigger computation. It is tempting because CDN edge caching genuinely reduces latency for static assets or cacheable web content served to distributed users.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

5 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?

easy
  • A.Cloud CDN
  • ✓ B.Cloud Memorystore
  • C.Cloud Spanner
  • D.BigQuery

Why B: Cloud Memorystore (B) is correct because it provides a managed in-memory cache (Redis or Memcached) that can store the results of identical prediction requests, reducing latency and load on the prediction endpoint. By caching responses keyed on the user profile or request parameters, subsequent identical requests can be served directly from Memorystore in microseconds, avoiding redundant model inference.

Variation 2. Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?

easy
  • ✓ A.Cloud Memorystore for Redis
  • B.Cloud CDN
  • C.Bigtable
  • D.Cloud Storage with object versioning

Why A: Cloud Memorystore for Redis is an in-memory data store that provides sub-millisecond latency, making it ideal for caching prediction responses. By caching identical prediction requests, you can reduce the number of calls to the Vertex AI endpoint, lowering latency and cost. Redis supports key-value storage with TTL, perfect for caching.

Variation 3. A company wants to cache predictions for identical requests to reduce latency and cost. They use Vertex AI Prediction with a custom container. Which GCP service should they use to implement prediction caching?

medium
  • A.Cloud Bigtable
  • ✓ B.Cloud Memorystore for Redis
  • C.Cloud Storage
  • D.Cloud Firestore

Why B: Cloud Memorystore for Redis is an in-memory data store with sub-millisecond latency, making it the ideal GCP service for caching prediction results keyed by request hash. It supports TTL-based expiration and high-throughput reads, which directly reduce latency and repeated model inference costs. This is the canonical GCP caching layer for Vertex AI prediction workloads.

Variation 4. Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?

hard
  • A.Switch to CPU-only instances to reduce cost.
  • B.Increase maxReplicas to handle the load without caching.
  • C.Set up a Cloud CDN in front of the endpoint.
  • ✓ D.Use Cloud Memorystore to cache prediction results.

Why D: Cloud Memorystore (Redis) in front of the Vertex AI endpoint lets you cache prediction results keyed by user/request signature, so repeated similar requests within short windows are served from cache instead of hitting the GPU-backed model. This reduces both latency (cache hit is sub-millisecond) and cost (fewer GPU inference calls).

Variation 5. A company deploys a model on Vertex AI Endpoints for real-time inference. They need to minimize latency for prediction requests that are identical to previous requests. Which approach should they use?

medium
  • A.Use a regional load balancer with session affinity.
  • ✓ B.Implement a caching layer using Cloud Memorystore with request hashing.
  • C.Use Cloud CDN to cache prediction responses.
  • D.Enable prediction caching on Vertex AI Endpoints.

Why B: Caching identical prediction requests using Cloud Memorystore with request hashing reduces latency by serving cached responses directly from an in-memory cache, avoiding redundant model inference. This approach is ideal for real-time inference where many requests are identical, as it bypasses the model endpoint entirely for cached requests, minimizing response time.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.