NCP-GENL Evaluation Practice Question
A team is evaluating a large language model for a question-answering system using NVIDIA NeMo Evaluation. They need to assess both the relevance of the answer to the question and its factual correctness. (Choose two.)
⚠ Common exam trap
The trap here is assuming that overlap-based metrics like F1 or BLEU capture both relevance and factual correctness, when they primarily measure lexical similarity and can be misled by paraphrasing or fluent hallucinations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Factual consistency score using an NLI model
The team must evaluate relevance and factual correctness. Answer relevance score directly measures how well the answer addresses the question, while factual consistency score via NLI checks whether the answer is factually supported. Exact Match, token F1, and BLEU focus on string overlap and do not separately assess relevance and factual correctness, making them less suitable for this dual requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
F1 score over tokens
Why it's wrong here
Token-level F1 measures the overlap between generated and reference answers, balancing precision and recall. While it captures partial correctness, it does not evaluate whether the answer is relevant to the question; an answer could overlap with the reference but be irrelevant if the reference is misaligned. Thus, it does not fully address both criteria.
- ✗
BLEU score
Why it's wrong here
BLEU measures n-gram precision against reference answers, primarily used in machine translation. It does not assess relevance to the question or factual correctness; a high BLEU score can be achieved with fluent but irrelevant or incorrect answers. Therefore, BLEU is not suitable for evaluating the two specified aspects.
- ✗
Exact Match (EM)
Why it's wrong here
Exact Match measures whether the generated answer exactly matches the reference answer after normalization. It is a strict metric that does not assess relevance to the question or factual correctness beyond exact string match. In many QA scenarios, answers can be phrased differently but still correct, so EM alone is insufficient for evaluating both relevance and correctness.
- ✓
Factual consistency score using an NLI model
Why this is correct
A factual consistency score uses an NLI model to check whether the generated answer is entailed by a trusted knowledge source (e.g., a reference document). This directly measures factual correctness, the second required criterion. It is robust to paraphrasing and can be automated within NeMo Evaluation to flag hallucinations or contradictions.
- ✓
Answer relevance score from a QA evaluation model
Why this is correct
An answer relevance score, often computed by a model like a cross-encoder or a specialized QA evaluator, assesses how well the answer addresses the question. This directly measures relevance, which is one of the two required criteria. It can be integrated into NeMo Evaluation to provide a quantitative relevance metric for each generated answer.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.