A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo to generate investment summaries. They need to evaluate the model's output for factual consistency against a set of gold-standard analyst reports. The team wants a metric that measures the longest common subsequence between the generated summary and the reference summary, ignoring word order. Which evaluation metric should they use?
Trap 1: BERTScore
BERTScore uses contextual embeddings to compute similarity between tokens, capturing semantic equivalence rather than surface-level n-gram overlap. It does not measure longest common subsequence; instead, it computes cosine similarities and aggregates them. While useful for semantic evaluation, it does not align with the requirement to measure LCS between generated and reference summaries.
Trap 2: METEOR
METEOR computes a harmonic mean of precision and recall based on unigram matches, incorporating synonyms and stemming. While it allows some word order flexibility, it does not explicitly compute the longest common subsequence. METEOR aligns words and computes chunk penalties, but the metric is not defined by LCS, so it does not meet the specific requirement.
Trap 3: BLEU
BLEU measures precision of n-grams, typically up to 4-grams, and includes a brevity penalty. It focuses on exact word order and does not use longest common subsequence, making it less appropriate for summarization where flexible ordering is acceptable. BLEU is more common in machine translation and would not satisfy the requirement of ignoring word order.
- A
ROUGE-L
ROUGE-L computes the longest common subsequence between the generated and reference summaries, which captures fluency and in-order matching without requiring contiguous n-grams. It is well-suited for summarization tasks where word order may vary but key content overlap matters, matching the requirement to ignore strict word order and measure factual consistency via LCS.
- B
BERTScore
Why it fails: BERTScore uses contextual embeddings to compute similarity between tokens, capturing semantic equivalence rather than surface-level n-gram overlap. It does not measure longest common subsequence; instead, it computes cosine similarities and aggregates them. While useful for semantic evaluation, it does not align with the requirement to measure LCS between generated and reference summaries.
- C
METEOR
Why it fails: METEOR computes a harmonic mean of precision and recall based on unigram matches, incorporating synonyms and stemming. While it allows some word order flexibility, it does not explicitly compute the longest common subsequence. METEOR aligns words and computes chunk penalties, but the metric is not defined by LCS, so it does not meet the specific requirement.
- D
BLEU
Why it fails: BLEU measures precision of n-grams, typically up to 4-grams, and includes a brevity penalty. It focuses on exact word order and does not use longest common subsequence, making it less appropriate for summarization where flexible ordering is acceptable. BLEU is more common in machine translation and would not satisfy the requirement of ignoring word order.