A developer is configuring a RAG application and needs to ensure that the LLM response is based on specific, trusted document snippets. Which technique, when implemented correctly, helps mitigate hallucination by grounding the response in provided context?
RAG grounds the model's output by providing relevant, factual information from trusted sources within the prompt. This context-based approach limits the model's tendency to hallucinate by forcing it to answer based on the provided document snippets, which are retrieved via similarity search before the generation step occurs.
Why this answer
Retrieval Augmented Generation (RAG) is the primary technique for grounding LLM responses. By retrieving relevant, trusted documents from a vector store based on a user's query and injecting those documents into the LLM's prompt, the developer forces the model to synthesize an answer based on specific retrieved context rather than relying solely on its internal training data. This significantly reduces hallucinations and increases the accuracy and relevance of the generated responses.
Exam trap
Candidates often confuse RAG with Fine-tuning or Prompt Engineering. They assume the model's internal weights are being updated, when RAG is strictly about providing external context at inference time.