Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?

⚠ Common exam trap

PMLE often tests the distinction between caching layers — candidates may pick Cloud CDN for 'caching' without realizing CDN only caches HTTP responses at the edge and cannot key on arbitrary prediction payloads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Memorystore

Cloud Memorystore provides a managed Redis or Memcached in-memory cache that can store frequent identical prediction requests and their responses, dramatically reducing latency and backend load. It is the standard Google Cloud service for application-level caching in front of Vertex AI endpoints.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Bigtable

    Why it's wrong here

    Cloud Bigtable is a wide-column NoSQL store for high-throughput analytical and time-series workloads, not a low-latency cache for identical prediction requests. Vertex AI prediction caching or Memorystore fills that role. Bigtable is correct for massive-scale key-range reads and writes, such as IoT or telemetry data.

  • ✗

    Cloud CDN

    Why it's wrong here

    Cloud CDN caches HTTP responses at edge locations for static or cacheable web content, not Vertex AI prediction payloads, and cannot key on request bodies. Vertex AI online prediction's own request-response caching or a dedicated cache layer handles identical predictions. Cloud CDN suits serving static assets and media to global users.

  • ✓

    Cloud Memorystore

    Why this is correct

    Cloud Memorystore offers a managed in-memory cache that the serving path can query before invoking the model, returning stored responses for repeated identical requests. This reduces endpoint invocations, lowering both latency and cost as the stem requires.

  • ✗

    Cloud SQL

    Why it's wrong here

    Cloud SQL is a managed relational database for transactional workloads; it stores and queries structured rows, not cached prediction responses keyed by request. Vertex AI prediction caching or Memorystore would serve that. Cloud SQL is correct when the requirement is a managed MySQL, PostgreSQL or SQL Server database for application data.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.