Reduce Azure AI Search Query Latency
Your Azure AI Search index is experiencing high query latency. You have enabled semantic search and custom scoring profiles. You need to reduce latency without degrading search quality. Which action should you take?
Quick Answer
The answer is to increase the number of replicas. Adding replicas distributes the query load across multiple identical copies of your index, enabling parallel processing of search requests which directly reduces query latency. This approach preserves search quality because it does not alter the underlying search logic, custom scoring profiles, or semantic enrichment configurations. On the AI-102 exam, this scenario tests your understanding of scaling strategies in Azure AI Search, often appearing as a distractor where candidates might mistakenly adjust the partition count or modify the index schema. A common trap is confusing replicas (for query performance) with partitions (for indexing throughput and storage). Remember the memory tip: “Replicas for reads, partitions for writes” to quickly recall that adding replicas is the correct move when latency is the issue.
⚠ Common exam trap
It's easy for candidates to confuse partitions (which affect storage and indexing speed) with replicas (which affect query throughput), leading them to incorrectly reduce partitions or disable features instead of scaling query capacity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of replicas.
Increasing the number of replicas distributes query load across multiple copies of the index, allowing parallel processing of search requests. This directly reduces query latency without altering the search logic, scoring profiles, or semantic enrichment, thus preserving search quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove custom scoring profiles.
Why it's wrong here
Removing custom scoring profiles changes how results are ranked, altering relevance rather than reducing query latency, so quality degrades. It is tempting because scoring profiles add processing, but they are not the latency source the scenario asks you to address.
- ✓
Increase the number of replicas.
Why this is correct
Query latency stems from limited compute, not index size. Replicas add query-processing capacity and enable load balancing across nodes, reducing latency while preserving semantic ranking and scoring profiles. Partition changes would affect storage and indexing, not query throughput.
- ✗
Reduce the number of partitions.
Why it's wrong here
Reducing partitions lowers the index's query throughput capacity, which typically increases latency and can force rebuilds. It is tempting because fewer resources sound cheaper, but partitions scale query performance; replica count, not partition count, addresses latency.
- ✗
Disable semantic search.
Why it's wrong here
Disabling semantic search removes the semantic reranking stage, which is what the stem requires to be preserved, so search quality drops. It is tempting because semantic reranking adds latency, but the requirement is to reduce latency without degrading quality.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on AI-102
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. You manage an Azure AI Search service that indexes legal documents. The search latency is high, and you need to improve query performance without reducing index size. Which action should you take?
hard- A.Upgrade to a higher pricing tier
- B.Increase the number of partitions
- C.Reduce the number of searchable fields
- ✓ D.Increase the number of replicas
Why D: Increasing the number of replicas distributes query load across multiple copies of the index, which directly improves query throughput and reduces latency. Replicas are designed for scaling query operations without changing the index size or storage capacity.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.