easyMultiple Choice
AIF-C01 Practice Question: Evaluate the quality of a text generation model…
A company wants to evaluate the quality of a text generation model for a summarization task. They have reference summaries written by humans. Which automated metric compares the generated summary to the reference by measuring n-gram overlap?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE
ROUGE (Recall-Oriented Understudy for Gisting Evaluation) measures n-gram overlap between generated and reference summaries. BLEU is for translation, BERTScore uses embeddings, and perplexity measures language model confidence.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
BLEU
Why it's wrong here
BLEU measures precision of n-gram overlap, penalising length via brevity penalty, but it does not account for recall of the reference content, so it under-rewards summaries that capture all key points. It is designed for machine translation, where precision-oriented scoring of short outputs against references is the standard evaluation approach.
- ✗
Perplexity
Why it's wrong here
Perplexity scores how fluently a language model predicts a token sequence, requiring no reference summary at all, so it cannot compare output against human references. It is tempting because it is a standard intrinsic metric for language models, and would be correct when judging a model's fluency or comparing trained checkpoints.
- ✓
ROUGE
Why this is correct
ROUGE measures n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for an automated metric with human reference summaries. Recall-oriented variants such as ROUGE-N and ROUGE-L quantify how much reference content the generated summary captures, making it the standard choice for summarization evaluation.
- ✗
BERTScore
Why it's wrong here
BERTScore embeds tokens and computes cosine similarity between contextual vectors, so it never counts n-gram overlap; it captures semantic equivalence instead. It is tempting because it handles paraphrasing and synonymy that BLEU penalises, making it the right pick when generated text differs in wording but matches reference meaning.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.