Courseiva
easyMultiple Choice

AIF-C01 Practice Question: Which metric is commonly used to evaluate the…

Which metric is commonly used to evaluate the quality of a text summarization model by comparing the generated summary with a reference summary, measuring the overlap of n-grams?

⚠ Common exam trap

The AWS AI Practitioner exam often tests the distinction between BLEU (precision-focused, for translation) and ROUGE (recall-focused, for summarization), and candidates mistakenly choose BLEU because both involve n-gram overlap, but BLEU is not the primary metric for summarization quality.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ROUGE

ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is the standard metric for evaluating text summarization by measuring the overlap of n-grams (e.g., ROUGE-1, ROUGE-2, ROUGE-L) between a generated summary and a reference summary. It focuses on recall, capturing how much of the reference content is preserved in the generated summary, making it ideal for summarization tasks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Perplexity

    Why it's wrong here

    Perplexity measures how well a language model predicts a token sequence, requiring no reference summary, so it cannot score overlap against one. It is tempting because it is a standard intrinsic metric for language generation quality, and it would be the right choice when assessing a model's fluency without gold references.

  • ✗

    BLEU

    Why it's wrong here

    BLEU measures n-gram precision against reference translations, penalising length mismatch, so it does not capture summary recall or content coverage. It is tempting because BLEU is a standard overlap metric, and it would be correct for evaluating machine translation output rather than summarisation.

  • ✓

    ROUGE

    Why this is correct

    ROUGE scores summarisation by counting overlapping n-grams between the generated summary and the reference summary, reporting recall-oriented precision. That direct n-gram comparison is exactly the overlap metric the question describes, unlike embedding-based or classification metrics.

  • ✗

    BERTScore

    Why it's wrong here

    BERTScore compares contextual embeddings rather than counting n-gram overlap, so it does not measure the surface matching the question describes. It is tempting because it evaluates semantic similarity and handles paraphrases well, making it a strong choice when summaries reword the reference rather than reuse its exact phrasing.

About these practice questions

This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.