PMLE Serving and Scaling Models Practice Question
A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Approximate nearest neighbor (ANN) index with ScaNN
For large datasets requiring low latency, an approximate nearest neighbor (ANN) index is appropriate. The Scann algorithm (ScaNN) is used by Vertex AI Vector Search for ANN.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Brute-force index
Why it's wrong here
Brute-force compares the query against every vector, giving exact results but latency and cost scale linearly with 10 million embeddings, missing the 10ms target. It is correct for small datasets or ground-truth benchmarking; Vertex AI Vector Search's approximate nearest-neighbour index is required here.
- ✗
Hash-based index
Why it's wrong here
Hash-based indexes bucket vectors by hash, which supports exact-match lookup but not nearest-neighbour ranking, so similarity search quality collapses. They are appropriate for duplicate detection or exact retrieval; approximate nearest-neighbour indexing is needed for low-latency similarity over 10 million embeddings.
- ✗
Tree-based index
Why it's wrong here
Tree-based indexes partition space hierarchically and degrade at high dimensionality, where traversal prunes poorly and latency rises beyond the 10ms target. They suit lower-dimensional or smaller datasets; Vertex AI Vector Search's scalable approximate nearest-neighbour index is required for 10 million high-dimensional embeddings.
- ✓
Approximate nearest neighbor (ANN) index with ScaNN
Why this is correct
ScaNN is Google's ANN algorithm optimised for high-dimensional vectors, using anisotropic vector quantisation to deliver sub-10ms latency and high throughput. It satisfies both stated constraints on 10 million embeddings, whereas exact or tree-based indexes cannot meet that latency at this scale.
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.