hardMultiple Choice
AIF-C01 Practice Question: A legal firm uses Amazon Bedrock to generate…
A legal firm uses Amazon Bedrock to generate contract summaries. They want to evaluate the quality of summaries against human-written reference summaries. The evaluation should capture both the overlap of n-grams and the semantic similarity. Which combination of automated metrics is MOST appropriate?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ROUGE and BERTScore
ROUGE measures n-gram overlap (recall-oriented) for summarization, while BERTScore uses contextual embeddings to capture semantic similarity. BLEU is for translation. Combining ROUGE and BERTScore gives both lexical and semantic evaluation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Exact match and F1 score
Why it's wrong here
Exact match requires the generated summary to match the reference string character for character, which abstractive summaries from Amazon Bedrock almost never achieve. F1 over tokens measures overlap but captures no semantic similarity. ROUGE and BERTScore together satisfy both stated requirements.
- ✗
ROUGE and BLEU
Why it's wrong here
BLEU measures n-gram precision against references, designed for machine translation, and ignores recall, so it misses content in the reference summary that the generated text omits. ROUGE handles n-gram overlap for summarisation, but neither metric captures semantic similarity, which BERTScore provides.
- ✓
ROUGE and BERTScore
Why this is correct
ROUGE measures n-gram overlap with the reference summaries, satisfying the lexical requirement, while BERTScore uses contextual embeddings to capture semantic similarity beyond exact wording. Together they cover both constraints in the stem, unlike metrics addressing only one dimension.
- ✗
BLEU and BERTScore
Why it's wrong here
BLEU computes n-gram precision for translation and ignores recall, so it penalises nothing when the summary omits reference content. BERTScore supplies the semantic element, but the n-gram half is mismatched to summarisation; ROUGE's recall-oriented overlap is the correct pairing.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.