Courseiva

PMLE Serving and Scaling Models Practice Question

You are using Vertex AI Vector Search for a product recommendation system. Your index is updated with new embeddings every hour. To minimize query latency while keeping the index fresh, what should you do?

⚠ Common exam trap

PMLE often tests the trade-off between index freshness and query latency, and candidates wrongly assume a full rebuild is required for updates, missing the streaming update capability.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use streaming updates to insert new embeddings into the deployed index.

Vertex AI Vector Search supports streaming updates that let you insert, update, or delete datapoints in a deployed index without rebuilding it, so queries stay low-latency while the index remains fresh. This is the intended mechanism for near-real-time freshness with minimal disruption. Batch rebuilds and redeployments cause downtime and latency spikes, which the question explicitly asks to avoid.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use streaming updates to insert new embeddings into the deployed index.

    Why this is correct

    Streaming updates let Vector Search mutate the deployed index in place as new embeddings arrive, so hourly refreshes avoid the rebuild-and-redeploy cycle that batch updates require. This keeps the index fresh without taking it offline, directly satisfying the low-latency, hourly-freshness constraint in the stem.

  • ✗

    Rebuild the entire index hourly as a batch job and redeploy it.

    Why it's wrong here

    Rebuilding the whole index hourly discards the existing serving index and forces a full redeploy, adding latency and downtime that streaming updates avoid. It is tempting because batch rebuilds are correct when the embedding model or index schema changes and incremental updates cannot apply.

  • ✗

    Create a new index each hour and use traffic splitting to gradually shift traffic.

    Why it's wrong here

    Vertex AI Vector Search supports streaming updates to an existing index, so hourly rebuilds and traffic splitting add operational overhead without improving freshness. It is tempting because traffic splitting safely validates a new index, which is right when changing index configuration such as distance measure or shard size.

  • ✗

    Use a brute-force index instead of ANN to ensure accuracy after updates.

    Why it's wrong here

    Brute-force search computes exact distances against every vector, so query latency scales linearly with index size and rises sharply as hourly embeddings accumulate. It suits small, static datasets where recall matters more than speed. Vertex AI Vector Search's ANN indexes already support streaming updates, giving freshness without that penalty.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.