Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

When evaluating an LLM for factual accuracy, which metric is most effective at detecting hallucinations compared to simple word-overlap metrics?

⚠ Common exam trap

Candidates frequently choose traditional lexical overlap metrics like BLEU or ROUGE, failing to recognize that they only measure token matching and miss semantic hallucinations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

NLI-based entailment scoring

Hallucinations in LLMs are semantic errors that word-overlap metrics like BLEU or ROUGE fail to capture, as they only measure token matching. Model-based evaluation, using frameworks like RAGAS or NLI (Natural Language Inference) models, assesses the truthfulness of generated content against reference contexts. This is critical for building trustworthy enterprise AI, where factual accuracy is paramount and simple text similarity does not suffice for assessing the validity of generated information.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    BLEU Score

    Why it's wrong here

    BLEU is designed for machine translation and measures n-gram overlap. It is notoriously bad for hallucination detection because it does not understand semantic meaning; a model can generate a completely hallucinated but grammatically correct sentence that receives a high BLEU score if it shares some words with the reference.

  • ✗

    ROUGE-L

    Why it's wrong here

    ROUGE measures the longest common subsequence between reference and generated text. Like BLEU, it is a surface-level metric that ignores factual accuracy and semantic consistency. It cannot distinguish between a truthful response and a hallucination that happens to share similar sentence structures or keywords with the ground truth.

  • ✓

    NLI-based entailment scoring

    Why this is correct

    NLI models are trained to determine if one sentence entails another. By checking if the generated output is entailed by the retrieved context, you can mathematically quantify the factual grounding of the model's response. This effectively identifies hallucinations by detecting when the model makes claims unsupported by the provided source material.

  • ✗

    Perplexity

    Why it's wrong here

    Perplexity measures how well the model predicts the next token in a sequence, essentially quantifying its confidence. A model can be very confident while hallucinating facts, making perplexity a measure of model fluency and language acquisition rather than a tool for verifying the factual truthfulness of its generated output.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.