A financial services firm is deploying a generative AI assistant that answers employee questions about internal policies. Compliance requires that every response cite the exact policy document and section used. The assistant currently relies only on the foundation model's pretrained knowledge and frequently invents policy details. Which technique should the team implement to ground responses in the firm's own documents and produce citations?
Retrieval Augmented Generation embeds source documents, retrieves the passages most semantically similar to the user's question, and passes them into the prompt so the model answers from that supplied context. Because the retrieved chunks carry document and section metadata, the assistant can cite the exact source. It also keeps answers current as policies change without retraining.
Why this answer
Grounding a model in proprietary content requires supplying that content at inference time rather than relying on pretrained weights. Retrieval Augmented Generation retrieves the most relevant document chunks and places them in the prompt, so the model's answer is conditioned on real policy text and can reference the source document and section. Sampling or length adjustments do not add knowledge.
Exam trap
The trap here is believing that fine-tuning on domain text guarantees accurate citations, when only retrieval of the actual source documents provides verifiable provenance.