easyMultiple Choice
AIF-C01 Practice Question: A developer needs to evaluate the quality of a…
A developer needs to evaluate the quality of a text summarization model by comparing its output to reference summaries. Which automated metric measures the overlap of n‑grams between the generated and reference summaries?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE
ROUGE measures n‑gram overlap and is commonly used for summarization. BLEU is for translation, BERTScore uses embeddings, and METEOR accounts for synonyms — but the question specifically asks about n‑gram overlap.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
ROUGE
Why this is correct
ROUGE computes recall-oriented n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for an automated metric measuring n-gram coincidence. Its variants (ROUGE-1, ROUGE-2, ROUGE-L) quantify unigram, bigram and longest-common-subsequence matching, making it the standard summarisation evaluation metric.
- ✗
BERTScore
Why it's wrong here
BERTScore embeds tokens and computes cosine similarity, capturing semantic equivalence rather than literal n-gram overlap, so it does not measure matching n-grams. It suits evaluating paraphrased or semantically varied outputs; ROUGE, which counts n-gram recall, is the metric described here.
- ✗
METEOR
Why it's wrong here
METEOR aligns unigrams with synonyms and stemming, then computes a harmonic mean of precision and recall with a fragmentation penalty; it does not simply measure raw n-gram overlap. It suits translation evaluation rewarding semantic matches, whereas ROUGE counts n-gram recall for summarisation.
- ✗
BLEU
Why it's wrong here
BLEU computes modified n-gram precision between candidate and reference text, but it is designed for machine translation evaluation, not summarisation. For summarisation, ROUGE measures n-gram recall against reference summaries, which is the overlap metric the question describes.
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.