A company is building an agent that uses Azure OpenAI to answer questions from a large document library. The agent must use a Retrieval Augmented Generation (RAG) pattern. Which TWO actions should the team take to implement RAG effectively?
RAG requires a searchable vector index so the agent can retrieve semantically similar chunks. Indexing documents into Azure Cognitive Search with embeddings satisfies the retrieval constraint, enabling grounded answers rather than relying solely on the model's parametric knowledge.
Why this answer
Option C is correct because RAG requires an external knowledge store that supports semantic similarity search, and indexing the documents into a vector database such as Azure Cognitive Search (or Azure AI Search) creates embeddings that let the agent retrieve the most relevant passages at query time. Option E is correct because the defining step of Retrieval Augmented Generation is retrieving the top-k relevant document chunks and injecting them into the prompt as grounding context before the model generates its answer, which keeps responses accurate and current without retraining. Option A is not appropriate because no model can memorize an entire large document library, and relying on memorization defeats the purpose of retrieval.
Option B is not appropriate because fine-tuning teaches style and task behavior rather than reliably storing and retrieving factual document content, and it is not the RAG mechanism. Option D is not appropriate because training a custom language model from scratch is prohibitively expensive and unnecessary when Azure OpenAI plus retrieval already solves the problem.
Exam trap
The trap here is that candidates often confuse fine-tuning (which adapts model behavior) with RAG (which augments prompts with retrieved data), leading them to select Option B instead of understanding that RAG requires an external retrieval step and vector index.