A team is tuning a large language model for a question-answering task. They notice the model gives high confidence scores to answers that are factually incorrect. Which evaluation metric should they primarily use to detect this overconfidence problem?
Trap 1: Perplexity
Perplexity measures how fluently a model predicts token sequences, not whether stated confidence matches factual accuracy, so it cannot expose overconfidence. It is tempting because perplexity is a standard intrinsic metric for language modelling and detects distributional drift; it would suit comparing model versions or detecting unfamiliar inputs, not calibration of answer correctness.
Trap 2: BLEU score
BLEU compares n-gram overlap between generated and reference text, so a confidently wrong answer sharing wording with a reference scores well and the miscalibration stays hidden. It is tempting because BLEU is a long-standing translation metric; it would be the right choice when scoring surface similarity to reference translations, not for measuring confidence against factual correctness.
Trap 3: ROUGE-L
ROUGE-L measures longest-common-subsequence overlap with a reference, which rewards fluent phrasing regardless of factual truth, leaving overconfident errors undetected. It is tempting because ROUGE is standard for summarisation evaluation; it would be correct when judging how much reference content a summary recalls, not for assessing whether confidence scores are calibrated.
- A
Perplexity
Why it fails: Perplexity measures how fluently a model predicts token sequences, not whether stated confidence matches factual accuracy, so it cannot expose overconfidence. It is tempting because perplexity is a standard intrinsic metric for language modelling and detects distributional drift; it would suit comparing model versions or detecting unfamiliar inputs, not calibration of answer correctness.
- B
Expected Calibration Error (ECE)
Expected Calibration Error directly measures the gap between predicted confidence and actual accuracy across probability bins, so systematically high confidence on wrong answers produces a large ECE value. This satisfies the stem's overconfidence constraint, unlike accuracy or F1, which ignore confidence entirely and cannot expose miscalibration.
- C
BLEU score
Why it fails: BLEU compares n-gram overlap between generated and reference text, so a confidently wrong answer sharing wording with a reference scores well and the miscalibration stays hidden. It is tempting because BLEU is a long-standing translation metric; it would be the right choice when scoring surface similarity to reference translations, not for measuring confidence against factual correctness.
- D
ROUGE-L
Why it fails: ROUGE-L measures longest-common-subsequence overlap with a reference, which rewards fluent phrasing regardless of factual truth, leaving overconfident errors undetected. It is tempting because ROUGE is standard for summarisation evaluation; it would be correct when judging how much reference content a summary recalls, not for assessing whether confidence scores are calibrated.