NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A machine learning engineer is evaluating a language model's performance on a text summarization task. The model achieves a BLEU score of 0.45 and a ROUGE-L score of 0.62 on the test set. The engineer wants to understand how well the model captures the overall meaning of the source documents. Which evaluation metric should they prioritize?
⚠ Common exam trap
The trap here is assuming that any high score indicates good summarization, when different metrics measure different aspects and only ROUGE-L directly evaluates recall of reference content.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE-L score, because it measures longest common subsequence and recall of reference content.
ROUGE-L is designed for summarization evaluation and measures recall of the longest common subsequence, which correlates with how much essential content from the reference is included. BLEU focuses on precision and is better for translation. Perplexity and accuracy do not assess summary quality against references. Therefore, ROUGE-L is the most appropriate metric for capturing overall meaning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Perplexity, because it measures how well the model predicts the next token.
Why it's wrong here
Perplexity measures language modeling confidence on a token level, not summarization quality. A model can have low perplexity but still produce summaries that omit critical details or hallucinate. Perplexity does not compare against reference summaries, so it cannot assess how well the overall meaning is captured.
- ✓
ROUGE-L score, because it measures longest common subsequence and recall of reference content.
Why this is correct
ROUGE-L evaluates the longest common subsequence between generated and reference summaries, emphasizing recall of important content. This aligns well with summarization goals, as it rewards including key information from the source. A high ROUGE-L indicates the model captures the essential meaning and structure of the reference summaries.
- ✗
Accuracy, because it measures the percentage of correctly predicted tokens.
Why it's wrong here
Token-level accuracy is not meaningful for free-text generation like summarization because there are many valid ways to phrase a summary. It does not account for semantic equivalence or recall of key points. Accuracy would penalize valid paraphrases and is not a standard metric for summarization quality.
- ✗
BLEU score, because it measures n-gram overlap with reference summaries.
Why it's wrong here
BLEU is primarily designed for machine translation and focuses on precision of n-grams. It does not account for recall or long-range coherence, making it less suitable for summarization where capturing the main ideas is key. A high BLEU score can be achieved with fluent but incomplete summaries, so it is not the best indicator of meaning preservation.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.