Courseiva
Evaluation →mediumMultiple Choice

NCP-GENL Evaluation Practice Question

A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?

⚠ Common exam trap

The trap here is assuming that high accuracy on standard benchmarks implies consistent behavior across paraphrased inputs, which it does not.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Semantic similarity between outputs for paraphrased inputs

The team needs to measure output consistency across semantically equivalent inputs. Semantic similarity between responses to paraphrased queries directly captures this variance, unlike lexical overlap metrics such as BLEU or ROUGE, which compare to references. Perplexity reflects language modeling confidence, not answer stability. Therefore, semantic similarity is the correct choice for quantifying inconsistency in the fine-tuned model's responses.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Semantic similarity between outputs for paraphrased inputs

    Why this is correct

    Semantic similarity measures how closely the model's responses align in meaning when the same question is asked with different wording. High similarity indicates consistent understanding and stable generation. In this scenario, computing similarity across paraphrased loan queries directly quantifies the observed inconsistency. This metric is appropriate for detecting and reducing output variance.

  • ✗

    ROUGE score

    Why it's wrong here

    ROUGE evaluates recall-oriented overlap with reference summaries, common in summarization. It cannot detect whether the model yields consistent answers to semantically equivalent queries. Using ROUGE here would measure similarity to a reference, not the variance across multiple phrasings of the same question. Therefore, it does not address the inconsistency issue.

  • ✗

    BLEU score

    Why it's wrong here

    BLEU measures n-gram overlap between generated and reference texts, typically used in translation. It does not assess response consistency across paraphrased inputs. In this scenario, BLEU would penalize valid paraphrases that differ lexically from a reference, failing to capture the semantic stability the team needs. Thus, it is unsuitable for quantifying inconsistency.

  • ✗

    Perplexity

    Why it's wrong here

    Perplexity quantifies how well a language model predicts a sample, reflecting fluency and confidence. It does not compare outputs for the same semantic query. A model could have low perplexity yet produce divergent answers to paraphrases. In this scenario, perplexity would not reveal the inconsistency the team observes, making it the wrong metric.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.