When evaluating an LLM for factual accuracy, which metric is most effective at detecting hallucinations compared to simple word-overlap metrics?
NLI models are trained to determine if one sentence entails another. By checking if the generated output is entailed by the retrieved context, you can mathematically quantify the factual grounding of the model's response. This effectively identifies hallucinations by detecting when the model makes claims unsupported by the provided source material.
Why this answer
Hallucinations in LLMs are semantic errors that word-overlap metrics like BLEU or ROUGE fail to capture, as they only measure token matching. Model-based evaluation, using frameworks like RAGAS or NLI (Natural Language Inference) models, assesses the truthfulness of generated content against reference contexts. This is critical for building trustworthy enterprise AI, where factual accuracy is paramount and simple text similarity does not suffice for assessing the validity of generated information.
Exam trap
Candidates frequently choose traditional lexical overlap metrics like BLEU or ROUGE, failing to recognize that they only measure token matching and miss semantic hallucinations.