NCP-GENL Evaluation Practice Question
A healthcare AI team is evaluating a large language model for clinical note summarization. They want to measure whether the generated summaries contain fabricated information not present in the source notes. Which evaluation approach is most appropriate for detecting hallucinated content?
⚠ Common exam trap
Watch out — candidates often confuse fluency metrics like perplexity with factual consistency metrics, when hallucination detection requires comparing the summary against the source for entailment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a natural language inference (NLI) model to check entailment between the source note and the generated summary.
Natural language inference (NLI) is effective for hallucination detection because it evaluates whether the generated summary is logically entailed by the source document. If the summary introduces unsupported information, the NLI model will not predict entailment. Other metrics like perplexity, BLEU, and ROUGE-L focus on fluency or surface overlap and cannot reliably identify fabricated content. Therefore, NLI-based checking is the most appropriate approach.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a natural language inference (NLI) model to check entailment between the source note and the generated summary.
Why this is correct
NLI models determine whether a hypothesis (the summary) is entailed by a premise (the source note). If the summary contains fabricated information, the NLI model would predict contradiction or neutral rather than entailment. This approach directly assesses factual consistency and is widely used for hallucination detection in summarization, making it the most appropriate method here.
- ✗
Compute the perplexity of the generated summaries on a held-out set of clinical notes.
Why it's wrong here
Perplexity measures how well a language model predicts a sample, reflecting fluency and likelihood. It does not compare the generated summary against the source note to detect fabricated content. A low perplexity indicates the model finds the text probable, but the text could still contain hallucinations. Thus, perplexity is not suitable for detecting factual inconsistencies or fabricated information.
- ✗
Calculate the BLEU score between the generated summary and a reference summary written by a clinician.
Why it's wrong here
BLEU compares n-gram overlap between generated and reference summaries. While it can indicate similarity, it does not verify factual consistency against the source note. A summary could have high BLEU by copying phrases from the reference but still include hallucinations not in the source. BLEU is not designed to detect fabricated content, so it is inadequate for this purpose.
- ✗
Measure the ROUGE-L score between the generated summary and the source note.
Why it's wrong here
ROUGE-L computes longest common subsequence between generated and reference texts, typically used for summarization quality. If the reference is the source note, it measures overlap but does not assess whether the summary is factually supported. A summary could share many sequences with the source yet still introduce fabricated details. ROUGE-L does not detect hallucinations.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.