NCP-GENL Evaluation Practice Question
A team is using NVIDIA NeMo Evaluator to assess a Llama-3-70B model fine-tuned for medical question answering. They want to evaluate both the correctness of answers and the model's ability to avoid hallucinating unsupported facts. Which two evaluation strategies should they implement? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any automated metric like BLEU or perplexity can substitute for factual verification, when hallucination detection requires external knowledge grounding.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a fact-checking module that cross-references generated statements with a trusted medical knowledge base.
To evaluate correctness, exact match and F1 score against gold answers provide a direct comparison. To detect hallucinations, a fact-checking module that verifies statements against a trusted medical knowledge base is essential. Together, these strategies cover both aspects. BLEU and perplexity do not measure factual accuracy, and a blanket guardrail against numbers is ineffective.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Compute BLEU score on the generated answers.
Why it's wrong here
BLEU score measures n-gram overlap and is designed for machine translation. In medical QA, correct answers may use different phrasing, so BLEU would unfairly penalize valid responses. It also does not assess factual correctness or hallucination. Thus, BLEU is not suitable for this evaluation goal.
- ✗
Calculate perplexity of the model on a held-out set of medical questions.
Why it's wrong here
Perplexity measures how well the model predicts the next token and is not a direct measure of factual correctness or hallucination. A model can have low perplexity yet still generate false statements. Therefore, perplexity alone does not satisfy the need to evaluate correctness or detect hallucinations.
- ✓
Use a fact-checking module that cross-references generated statements with a trusted medical knowledge base.
Why this is correct
A fact-checking module can verify each claim in the model's output against a curated medical knowledge base, directly detecting hallucinations or unsupported facts. This approach provides a granular, evidence-based assessment of factual accuracy, which is critical in healthcare. Integrating such a module with NeMo Evaluator allows automated flagging of unsupported statements.
- ✓
Compare model outputs against gold-standard answers using exact match and F1 score.
Why this is correct
Exact match and F1 score provide a quantitative measure of answer correctness relative to ground truth. In medical QA, where precise terminology matters, these metrics can indicate how often the model produces the expected answer. While they may not capture semantic equivalence perfectly, they are standard and efficient for large-scale evaluation, making them a valid component of the assessment.
- ✗
Run the model through NeMo Guardrails with a policy that blocks any output containing numbers.
Why it's wrong here
Blocking outputs with numbers is an arbitrary rule that would not effectively detect hallucinations; many correct medical answers include numbers (e.g., dosages). This approach would introduce false positives and fail to assess factual accuracy. It does not address the core evaluation goals of correctness and hallucination detection.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.