AIF-C01 Fundamentals of Generative AI Practice Question
A data scientist is evaluating foundation models for a text summarization task and wants to use a standard metric. Which metric is commonly used to assess the quality of generated summaries?
⚠ Common exam trap
Candidates often confuse BLEU and ROUGE, mistakenly selecting BLEU for summarization because it is a well-known metric for text generation. However, BLEU is designed for machine translation precision, while ROUGE focuses on recall and n-gram overlap, making it the standard for summarization evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE
ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is the standard metric for text summarization. It measures the overlap of n-grams, word sequences, or word pairs between the generated summary and reference summaries, focusing on recall. This makes it particularly suitable for evaluating how well the generated summary captures the key content of the reference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
F1 score
Why it's wrong here
F1 score balances precision and recall for classification, comparing predicted labels against gold labels. Summarisation output is free text with no discrete class labels, so precision and recall cannot be computed directly. It is tempting because F1 is standard for NLP classification tasks, but summarisation quality is measured with overlap metrics such as ROUGE.
- ✓
ROUGE
Why this is correct
ROUGE measures n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for a standard summarisation metric. Recall-oriented variants such as ROUGE-N and ROUGE-L capture how much reference content the model reproduced, making it the conventional benchmark for evaluating summary quality.
- ✗
BLEU
Why it's wrong here
BLEU measures n-gram precision against reference translations, penalising short output, and was designed for machine translation. Summarisation evaluation instead uses recall-oriented overlap metrics such as ROUGE, which capture how much reference content the summary retains. BLEU is tempting because both tasks generate text, but its precision weighting suits translation, not summarisation.
- ✗
Accuracy
Why it's wrong here
Accuracy requires discrete predicted labels matched against ground-truth labels. Summarisation produces free-form text with many valid phrasings, so exact-match accuracy is undefined and near-zero. It is tempting because accuracy is the default metric for classification, but generated summaries are scored with overlap metrics such as ROUGE instead.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.