NCP-GENL Evaluation Practice Question
A team is evaluating a retrieval-augmented generation (RAG) pipeline using NVIDIA NeMo Evaluation. They notice that the generated answers are fluent but sometimes contradict the retrieved documents. Which evaluation approach best identifies this issue?
⚠ Common exam trap
The trap here is relying on reference-based metrics like BLEU when the issue is faithfulness to retrieved context, not similarity to a reference answer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a natural language inference (NLI) model to check entailment between retrieved documents and generated answers.
The problem is factual inconsistency between retrieved documents and generated answers. Natural language inference directly evaluates entailment, identifying contradictions. Other metrics like BLEU, perplexity, or document length do not measure this relationship, making NLI the appropriate choice for this RAG evaluation scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Compute BLEU score against reference answers.
Why it's wrong here
BLEU compares n-gram overlap with reference answers, but references may not exist or may not reflect retrieved documents. Contradictions with retrieved context are not captured by BLEU, as it focuses on surface similarity to references, not factual consistency with source documents. Thus, it fails to detect the contradiction problem.
- ✗
Calculate perplexity of the generated answers.
Why it's wrong here
Perplexity measures language model confidence and fluency, not factual consistency with retrieved documents. A fluent but contradictory answer can have low perplexity. Therefore, perplexity cannot identify contradictions; it only indicates how predictable the text is under the model, which is unrelated to the retrieval grounding issue.
- ✗
Measure the average length of retrieved documents.
Why it's wrong here
Average document length is a retrieval statistic and does not assess the generated answer's fidelity to the documents. Contradictions can occur regardless of document length. Thus, this metric is irrelevant for detecting whether the generated answer contradicts the retrieved content, and it provides no insight into factual consistency.
- ✓
Use a natural language inference (NLI) model to check entailment between retrieved documents and generated answers.
Why this is correct
NLI models determine whether a hypothesis (generated answer) is entailed by, neutral to, or contradicts a premise (retrieved document). By running an NLI model, you can flag contradictions directly. This approach is robust for RAG evaluation because it assesses factual consistency without requiring reference answers, aligning with the goal of detecting contradictions.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.