NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A machine learning engineer is evaluating a generative LLM for a customer-facing question-answering system. The model produces fluent answers, but during testing it confidently states incorrect facts about company policies. The team wants a metric that specifically measures whether the model's output is supported by the provided source documents. Which evaluation approach is most appropriate?
⚠ Common exam trap
The trap here is choosing a familiar text-generation metric like BLEU or ROUGE because it is easy to compute, when the scenario specifically requires verifying factual support against source documents.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Faithfulness or groundedness evaluation that checks whether claims in the answer are entailed by the retrieved source passages
The team needs to know whether generated answers are supported by the source documents, which is precisely what faithfulness or groundedness evaluation measures. It checks entailment between answer claims and retrieved passages, catching confident hallucinations. Perplexity, BLEU, and ROUGE assess fluency or surface overlap with references, none of which guarantee that a stated policy fact is actually present in the source material.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Perplexity measured on a held-out set of company documents
Why it's wrong here
Perplexity measures how well the model predicts the next token in a text corpus. A low perplexity means the model finds the documents statistically likely, but it says nothing about whether a generated answer is faithful to a specific source. A model can be fluent and low-perplexity while still hallucinating policy details, so this metric does not address the requirement.
- ✗
ROUGE score computed against a large corpus of historical customer emails
Why it's wrong here
ROUGE measures recall-oriented n-gram overlap with reference texts. Historical customer emails are not authoritative policy sources and may contain outdated or incorrect information. A high ROUGE score would only indicate similarity to past emails, not factual correctness. This approach also requires reference summaries for every answer, which the team does not have for arbitrary policy questions.
- ✗
BLEU score computed between the model's answers and reference answers
Why it's wrong here
BLEU measures n-gram overlap with one or more reference answers. It rewards surface similarity, not factual support, and it penalizes correct paraphrases. In a policy QA setting, a factually correct answer phrased differently from the reference could score poorly, while a fluent but incorrect answer that copies reference wording could score well. It is not a faithfulness metric.
- ✓
Faithfulness or groundedness evaluation that checks whether claims in the answer are entailed by the retrieved source passages
Why this is correct
Faithfulness or groundedness metrics compare each claim in the generated answer against the provided source documents to determine whether the source supports it. This directly detects hallucinated policy details, which is the team's concern. It can be implemented with human annotation or with an entailment model, and it aligns with the requirement to measure support by source documents.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.