NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A machine learning engineer is evaluating a fine-tuned LLM for a customer-facing summarization task. The model produces fluent summaries, but the team needs to detect when the model generates content that is not supported by the source document. Which TWO evaluation approaches are appropriate for measuring factual consistency between the generated summary and the source? (Choose two.)
⚠ Common exam trap
The trap here is choosing reference-overlap metrics like BLEU or ROUGE because they are familiar, even though they do not verify that generated content is grounded in the source document.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Have human annotators label each summary sentence as supported or not supported by the source document.
Factual consistency requires checking whether the summary's claims are supported by the source document. Entailment-based metrics like FactCC or NLI models directly test this relationship, and human annotation provides a reliable ground truth. Metrics like BLEU, ROUGE, or perplexity measure overlap or fluency and can be fooled by fluent hallucinations, so they are not appropriate for this specific goal.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Compute BLEU score between the generated summary and a single human-written reference summary.
Why it's wrong here
BLEU measures n-gram overlap with a reference and is designed for translation. It does not check whether claims in the summary are supported by the source document, and a fluent but fabricated summary can still score well if it shares phrasing with the reference. For factual consistency, reference-based overlap metrics are insufficient.
- ✗
Calculate ROUGE-L recall against the source document.
Why it's wrong here
ROUGE-L measures longest common subsequence overlap, typically between a generated summary and a reference summary. Using it against the source document rewards copying source phrases but does not verify that the summary's claims are logically supported. A summary that copies unrelated source sentences could score high while being factually inconsistent.
- ✓
Have human annotators label each summary sentence as supported or not supported by the source document.
Why this is correct
Human annotation directly assesses whether each claim in the summary is grounded in the source. Annotators can catch subtle fabrications that automated metrics miss. While more expensive, it provides a reliable gold standard for factual consistency and is often used to validate automated metrics in production settings.
- ✗
Measure perplexity of the generated summary under the fine-tuned model.
Why it's wrong here
Perplexity measures how well the model predicts the summary text, reflecting fluency and likelihood under the model. A hallucinated summary can have low perplexity if the model is confident about its fabrication. Perplexity does not compare the summary to the source document, so it cannot detect unsupported content.
- ✓
Use an entailment-based metric such as FactCC or a natural language inference model to check whether the summary is entailed by the source.
Why this is correct
Entailment-based metrics treat the source document as the premise and the summary as the hypothesis, checking whether the summary's claims are logically supported. This directly targets factual consistency because unsupported statements fail the entailment test. Tools like FactCC or NLI models are standard for this purpose and do not require a human reference summary.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.