easyMultiple Choice
AIF-C01 Practice Question: Which metric is commonly used to evaluate the…
Which metric is commonly used to evaluate the quality of a text summarization model by comparing the generated summary with a reference summary, measuring the overlap of n-grams?
⚠ Common exam trap
The AWS AI Practitioner exam often tests the distinction between BLEU (precision-focused, for translation) and ROUGE (recall-focused, for summarization), and candidates mistakenly choose BLEU because both involve n-gram overlap, but BLEU is not the primary metric for summarization quality.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE
ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is the standard metric for evaluating text summarization by measuring the overlap of n-grams (e.g., ROUGE-1, ROUGE-2, ROUGE-L) between a generated summary and a reference summary. It focuses on recall, capturing how much of the reference content is preserved in the generated summary, making it ideal for summarization tasks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how well a language model predicts a token sequence, requiring no reference summary, so it cannot score overlap against one. It is tempting because it is a standard intrinsic metric for language generation quality, and it would be the right choice when assessing a model's fluency without gold references.
- ✗
BLEU
Why it's wrong here
BLEU measures n-gram precision against reference translations, penalising length mismatch, so it does not capture summary recall or content coverage. It is tempting because BLEU is a standard overlap metric, and it would be correct for evaluating machine translation output rather than summarisation.
- ✓
ROUGE
Why this is correct
ROUGE scores summarisation by counting overlapping n-grams between the generated summary and the reference summary, reporting recall-oriented precision. That direct n-gram comparison is exactly the overlap metric the question describes, unlike embedding-based or classification metrics.
- ✗
BERTScore
Why it's wrong here
BERTScore compares contextual embeddings rather than counting n-gram overlap, so it does not measure the surface matching the question describes. It is tempting because it evaluates semantic similarity and handles paraphrases well, making it a strong choice when summaries reword the reference rather than reuse its exact phrasing.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.