NCA-GENL Core Machine Learning and AI Knowledge Practice Question
When evaluating an LLM for factual accuracy, which metric is most effective at detecting hallucinations compared to simple word-overlap metrics?
⚠ Common exam trap
Candidates frequently choose traditional lexical overlap metrics like BLEU or ROUGE, failing to recognize that they only measure token matching and miss semantic hallucinations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NLI-based entailment scoring
Hallucinations in LLMs are semantic errors that word-overlap metrics like BLEU or ROUGE fail to capture, as they only measure token matching. Model-based evaluation, using frameworks like RAGAS or NLI (Natural Language Inference) models, assesses the truthfulness of generated content against reference contexts. This is critical for building trustworthy enterprise AI, where factual accuracy is paramount and simple text similarity does not suffice for assessing the validity of generated information.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
BLEU Score
Why it's wrong here
BLEU is designed for machine translation and measures n-gram overlap. It is notoriously bad for hallucination detection because it does not understand semantic meaning; a model can generate a completely hallucinated but grammatically correct sentence that receives a high BLEU score if it shares some words with the reference.
- ✗
ROUGE-L
Why it's wrong here
ROUGE measures the longest common subsequence between reference and generated text. Like BLEU, it is a surface-level metric that ignores factual accuracy and semantic consistency. It cannot distinguish between a truthful response and a hallucination that happens to share similar sentence structures or keywords with the ground truth.
- ✓
NLI-based entailment scoring
Why this is correct
NLI models are trained to determine if one sentence entails another. By checking if the generated output is entailed by the retrieved context, you can mathematically quantify the factual grounding of the model's response. This effectively identifies hallucinations by detecting when the model makes claims unsupported by the provided source material.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how well the model predicts the next token in a sequence, essentially quantifying its confidence. A model can be very confident while hallucinating facts, making perplexity a measure of model fluency and language acquisition rather than a tool for verifying the factual truthfulness of its generated output.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.