NCP-GENL Evaluation Practice Question
A financial services company has fine-tuned a Llama-3-70B model using NVIDIA NeMo for automated loan risk assessment. During evaluation, they observe that the model achieves high scores on standard accuracy metrics but produces inconsistent responses when the same query is phrased slightly differently. They need a metric that quantifies this inconsistency. Which evaluation metric should they use?
⚠ Common exam trap
The trap here is assuming that high accuracy on standard benchmarks implies consistent behavior across paraphrased inputs, which it does not.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Semantic similarity between outputs for paraphrased inputs
The team needs to measure output consistency across semantically equivalent inputs. Semantic similarity between responses to paraphrased queries directly captures this variance, unlike lexical overlap metrics such as BLEU or ROUGE, which compare to references. Perplexity reflects language modeling confidence, not answer stability. Therefore, semantic similarity is the correct choice for quantifying inconsistency in the fine-tuned model's responses.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Semantic similarity between outputs for paraphrased inputs
Why this is correct
Semantic similarity measures how closely the model's responses align in meaning when the same question is asked with different wording. High similarity indicates consistent understanding and stable generation. In this scenario, computing similarity across paraphrased loan queries directly quantifies the observed inconsistency. This metric is appropriate for detecting and reducing output variance.
- ✗
ROUGE score
Why it's wrong here
ROUGE evaluates recall-oriented overlap with reference summaries, common in summarization. It cannot detect whether the model yields consistent answers to semantically equivalent queries. Using ROUGE here would measure similarity to a reference, not the variance across multiple phrasings of the same question. Therefore, it does not address the inconsistency issue.
- ✗
BLEU score
Why it's wrong here
BLEU measures n-gram overlap between generated and reference texts, typically used in translation. It does not assess response consistency across paraphrased inputs. In this scenario, BLEU would penalize valid paraphrases that differ lexically from a reference, failing to capture the semantic stability the team needs. Thus, it is unsuitable for quantifying inconsistency.
- ✗
Perplexity
Why it's wrong here
Perplexity quantifies how well a language model predicts a sample, reflecting fluency and confidence. It does not compare outputs for the same semantic query. A model could have low perplexity yet produce divergent answers to paraphrases. In this scenario, perplexity would not reveal the inconsistency the team observes, making it the wrong metric.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.