NCP-GENL Evaluation Practice Question
A healthcare AI team is evaluating a fine-tuned GPT-based model for clinical note summarization using NVIDIA NeMo. They need to assess whether the model's summaries contain no fabricated medical facts. Which evaluation approach is most appropriate?
⚠ Common exam trap
Candidates often confuse semantic similarity to a reference summary with factual consistency against the source document, which are fundamentally different checks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a factual consistency metric such as SummaC or Q2
The critical requirement is detecting fabricated medical facts, which demands checking entailment between the summary and the source clinical note. Factual consistency metrics such as SummaC or Q2 are specifically designed for this purpose, unlike similarity-based metrics (BERTScore, ROUGE-L) that compare to references and may miss hallucinations. Perplexity assesses fluency, not factuality. Therefore, a factual consistency metric is the correct evaluation approach.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a factual consistency metric such as SummaC or Q2
Why this is correct
Factual consistency metrics like SummaC or Q2 evaluate whether generated summaries are entailed by the source document, detecting hallucinations. They are designed to catch fabricated facts by comparing summary claims against source content. In this healthcare scenario, applying such a metric directly addresses the need to ensure no invented medical facts, making it the most suitable approach.
- ✗
Measure BERTScore against reference summaries
Why it's wrong here
BERTScore computes semantic similarity between generated and reference texts using contextual embeddings. While it captures semantic overlap, it does not verify factual accuracy against source documents. A summary could be semantically similar to a reference yet contain fabricated facts not present in the source. In this clinical scenario, BERTScore alone would not detect hallucinations, so it is insufficient.
- ✗
Compute perplexity of the summaries
Why it's wrong here
Perplexity measures how well the language model predicts the summary text, reflecting fluency. It does not compare the summary to the source document for factual accuracy. A fluent but hallucinated summary could have low perplexity. In clinical note summarization, perplexity would not flag fabricated facts, so it is not a valid evaluation method for this requirement.
- ✗
Calculate ROUGE-L scores
Why it's wrong here
ROUGE-L measures longest common subsequence overlap with references, focusing on recall. It is commonly used in summarization but does not check factual consistency with the source. A model could score high by copying phrases while introducing false statements. For clinical notes, where fabricated facts are dangerous, ROUGE-L fails to provide the necessary factual verification, making it inappropriate.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.