A financial services company is fine-tuning an LLM to answer questions about internal policies. The base model performs well on general text but frequently invents policy numbers and effective dates. The team has a curated dataset of 5,000 question-answer pairs with correct citations. Which fine-tuning approach best addresses the hallucination of policy numbers and dates?
Supervised fine-tuning on curated question-answer pairs with correct citations directly teaches the model the desired input-output behavior. By training on examples that include accurate policy numbers and effective dates, the model learns to associate questions with grounded answers rather than plausible-sounding fabrications, which is the most direct remedy for the observed hallucination pattern.
Why this answer
Supervised fine-tuning on curated question-answer pairs with correct citations is the most direct way to teach the model the desired behavior. The training examples pair each question with grounded policy numbers and effective dates, so the model learns to reproduce accurate citations instead of inventing them. Continued pre-training, length-based RLHF, and higher inference temperature do not target the specific hallucination of policy details.
Exam trap
The trap here is reaching for continued pre-training or RLHF when the scenario already supplies labeled question-answer pairs, which are the natural input for supervised fine-tuning.