easyMultiple Select
AIF-C01 Practice Question: Using Amazon Bedrock with a foundation model for…
A company is using Amazon Bedrock with a foundation model for a text summarization task. They want to evaluate the quality of the summaries. Which TWO metrics are appropriate for evaluating the quality of generated summaries? (Select TWO)
⚠ Common exam trap
AWS often tests the distinction between metrics designed for generation tasks (ROUGE, BLEU) versus classification metrics (Accuracy) or model performance metrics (Perplexity, Latency), leading candidates to mistakenly select Accuracy or Perplexity for summarization evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE score
ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is correct because it measures n-gram and longest-common-subsequence overlap between generated summaries and reference summaries, which directly captures summary content coverage and is the standard metric for summarization quality. BLEU is also correct because it computes precision-based n-gram overlap between generated and reference text, and while traditionally used for machine translation, it is widely applied to summarization evaluation to assess how much of the generated summary matches the reference phrasing. Accuracy is not appropriate because summarization is a free-form generation task without a single discrete correct label, so classification-style accuracy cannot be computed meaningfully. Latency measures inference speed, not output quality, so it does not evaluate the summary's content. Perplexity measures how well a language model predicts a token sequence and reflects fluency/likelihood, but it does not compare generated summaries against references and thus is not a direct quality metric for summarization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
ROUGE score
Why this is correct
ROUGE measures n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for summary-quality evaluation. Recall-oriented variants capture whether key content from the source appears in the output, making it the standard metric for summarisation tasks rather than classification or retrieval.
- ✓
BLEU score
Why this is correct
BLEU compares n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for a quality metric suited to text summarisation. It quantifies how much of the reference wording the model reproduces, making it a standard, appropriate choice for evaluating generated summary fidelity against ground-truth text.
- ✗
Accuracy
Why it's wrong here
Accuracy measures classification correctness against labelled categories, which summarisation does not produce; summaries need overlap or similarity scoring instead. Accuracy is appropriate for tasks such as sentiment or intent classification, where each output maps to a discrete correct label.
- ✗
Latency
Why it's wrong here
Latency measures response time, not summary quality, so it cannot assess faithfulness or relevance of generated text. It is tempting because latency is a genuine operational metric for production inference, and would be the right measure when the requirement is throughput or response-time benchmarking rather than evaluating output quality.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how well a language model predicts a token sequence, not whether a summary is faithful or coherent. It is tempting because perplexity is a standard intrinsic metric for language models, and would suit comparing model checkpoints during pre-training rather than judging generated summary quality.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.