PMLE Serving and Scaling Models Practice Question
A company uses Vertex AI Vector Search for similarity search. They have a dataset of 10 million 512-dimensional vectors. Which index type should they choose for lowest latency at high recall?
⚠ Common exam trap
It's easy for candidates to assume brute-force is the only way to guarantee high recall, but the question explicitly asks for lowest latency at high recall, which is the exact trade-off that ANN indexes like ScaNN are designed to optimize.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Approximate nearest neighbor (ANN) index with Scann
For a dataset of 10 million 512-dimensional vectors, a brute-force (flat) index would be far too slow for low-latency queries. Approximate Nearest Neighbor (ANN) with ScaNN (Scalable Nearest Neighbors) is specifically designed by Google for high-dimensional vector search, offering sub-linear query time while maintaining high recall through techniques like anisotropic quantization and tree-based partitioning. This makes it the optimal choice for balancing latency and recall at this scale.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Brute-force (flat) index
Why it's wrong here
A brute-force flat index compares the query against every vector, giving exact recall but latency that scales linearly and becomes unacceptable at 10 million 512-dimensional vectors. It suits small datasets or ground-truth benchmarking; the low-latency, high-recall requirement here calls for a graph-based index such as ScaNN or HNSW.
- ✓
Approximate nearest neighbor (ANN) index with Scann
Why this is correct
Scann's approximate nearest neighbour index trades exact search for graph-based traversal, delivering far lower query latency at high recall on large vector sets. For 10 million 512-dimensional vectors, this satisfies the lowest-latency-at-high-recall constraint better than brute-force exact search.
- ✗
Tree-based index
Why it's wrong here
Tree-based indexes such as Annoy partition space hierarchically and degrade in recall at 512 dimensions, where distance concentration erodes their pruning. They suit low-dimensional or memory-constrained approximate search; Vertex AI Vector Search's low-latency, high-recall requirement at this scale calls for a graph-based index such as ScaNN or HNSW.
- ✗
Hashing-based index
Why it's wrong here
Hashing-based indexes compress vectors into buckets, trading recall for speed and memory, so they cannot sustain high recall across 10 million 512-dimensional vectors. They suit approximate deduplication or very large-scale candidate generation; Vertex AI Vector Search's low-latency, high-recall requirement calls for a graph-based index such as ScaNN or HNSW.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
3 more ways this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. You need to query a Vertex AI Vector Search index for nearest neighbours. The index is deployed on an endpoint. Which API method should you use to perform the query?
medium- ✓ A.projects.locations.indexEndpoints.findNeighbors
- B.projects.locations.indexes.match
- C.projects.locations.indexes.query
- D.projects.locations.endpoints.predict
Why A: The correct API method to query a deployed Vertex AI Vector Search index for nearest neighbors is `projects.locations.indexEndpoints.findNeighbors`. This method is specifically designed for vector similarity search against an index endpoint, returning the nearest neighbors for a given query vector. The other options either target the wrong resource (indexes instead of indexEndpoints) or use methods intended for different purposes like model prediction.
Variation 2. You are using Vertex AI Vector Search with an approximate nearest neighbor index. You need to update the index with new data every hour. The updates must be available for queries immediately. Which update method should you use?
medium- A.Recreate the index every hour using a scheduled job.
- B.Batch update by creating a new index and deploying it.
- ✓ C.Streaming updates using the streaming API.
- D.Use a brute-force index that supports real-time updates.
Why C: Vertex AI Vector Search supports streaming updates via its streaming API, which allows you to insert, update, or delete vectors in real time. This ensures that new data is immediately available for approximate nearest neighbor (ANN) queries without requiring index recreation or redeployment, meeting the requirement for hourly updates with instant query availability.
Variation 3. A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?
hard- A.Brute-force index
- B.Hash-based index
- C.Tree-based index
- ✓ D.Approximate nearest neighbor (ANN) index with ScaNN
Why D: For large datasets requiring low latency, an approximate nearest neighbor (ANN) index is appropriate. The Scann algorithm (ScaNN) is used by Vertex AI Vector Search for ANN.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.