Courseiva
Serving and Scaling Models →mediumMultiple Select

PMLE Serving and Scaling Models Practice Question

A company uses Vertex AI Matching Engine for real-time recommendations. They need to serve queries with low latency and support frequent updates. Which two configurations are appropriate? (Choose 2)

⚠ Common exam trap

The trap here is that in Google's Vertex AI Matching Engine, candidates often confuse batch updates with streaming updates, assuming that batch updates can be made frequent enough to approximate real-time, but they fail to recognize that batch updates require full index rebuilds, which introduce significant latency and downtime for serving.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable streaming updates for the index

Option B is correct because Vertex AI Matching Engine supports streaming updates, which allow the index to be modified incrementally as new data arrives without requiring a full rebuild, directly satisfying the requirement for frequent updates in a real-time recommendation system. Option D is correct because deploying the index to a Vertex AI Matching Engine endpoint is the standard mechanism for serving low-latency, real-time nearest-neighbor queries at scale, which is exactly what the scenario demands. Option A is not appropriate because storing the index in Cloud Storage and querying it via Python does not provide the managed, low-latency serving infrastructure of Matching Engine. Option C is not appropriate because a brute-force index performs exhaustive comparisons, which is far too slow for low-latency real-time serving at scale. Option E is not appropriate because batch-only updates cannot keep pace with frequent updates and would introduce staleness in a real-time recommendation system.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Store the index in Cloud Storage and query via Python

    Why it's wrong here

    Cloud Storage holds objects for batch ingestion, not low-latency vector search; querying it via Python adds retrieval latency and cannot serve online recommendations. It is tempting because it is cheap and familiar for storing embeddings; it would be correct as the source location for building or rebuilding an index.

  • ✓

    Enable streaming updates for the index

    Why this is correct

    Streaming updates allow new embeddings to be inserted or removed incrementally while the index remains queryable, avoiding the full rebuild and redeploy that batch updates force. This satisfies the frequent-update requirement without interrupting the low-latency online serving the stem specifies.

  • ✗

    Use a brute-force index for exact results

    Why it's wrong here

    Brute-force search compares the query against every vector, so latency scales with index size and cannot meet the low-latency serving requirement. It is tempting because it returns exact nearest neighbours; it would be correct for small datasets or accuracy benchmarking, not large-scale real-time recommendation.

  • ✓

    Deploy the index to a Vertex AI Matching Engine endpoint

    Why this is correct

    Deploying the index to a Matching Engine endpoint exposes it for online queries over the network, returning approximate nearest-neighbour results in milliseconds. This satisfies the low-latency serving requirement, since an undeployed index cannot answer real-time recommendation requests.

  • ✗

    Use batch updates only

    Why it's wrong here

    Batch updates leave the index stale between rebuilds, so newly added items are not served, failing the frequent-update requirement. It is tempting because batch rebuilds are operationally simple; they would be correct where the catalogue changes rarely and near-real-time freshness is unnecessary.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.