Courseiva
mediumMultiple Choice

AIF-C01 Practice Question: Evaluating the performance of a summarization…

A company is evaluating the performance of a summarization model using Amazon Bedrock model evaluation. They want an automated metric that measures how well the generated summary captures the meaning of the reference summary. Which metric is MOST suitable?

⚠ Common exam trap

The trap is that candidates often default to ROUGE-L because it is commonly used in summarization tasks, but for Amazon Bedrock's automated evaluation, BERTScore is preferred as it captures semantic meaning using contextual embeddings from BERT.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

BERTScore

BERTScore is the most suitable metric because it uses contextual embeddings from BERT to compute semantic similarity between generated and reference summaries, capturing meaning rather than exact n-gram overlap. This makes it ideal for summarization evaluation where paraphrasing and semantic equivalence are critical.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    BLEU

    Why it's wrong here

    BLEU measures n-gram precision against reference text, designed for machine translation, and ignores recall, so it misses whether the summary covers the reference's meaning. It is tempting because BLEU is a widely recognised automated metric, and it would be correct when scoring translation output against reference translations.

  • ✗

    ROUGE-L

    Why it's wrong here

    ROUGE-L scores longest common subsequence overlap with the reference, capturing surface wording rather than meaning, so paraphrased summaries score poorly. It is tempting because ROUGE is the standard summarisation metric, and ROUGE-L would be the right choice when lexical overlap with a reference summary is the acceptance criterion.

  • ✗

    Accuracy

    Why it's wrong here

    Accuracy requires discrete labelled outcomes, but summarisation produces free text with no single correct label, so it cannot measure meaning capture. It is tempting because accuracy is the default metric for classification tasks, and it would be correct when evaluating a model that assigns each input to one predefined class.

  • ✓

    BERTScore

    Why this is correct

    BERTScore computes token-level semantic similarity using contextual embeddings, so it rewards summaries that preserve meaning even when wording differs from the reference. That directly satisfies the requirement to measure how well generated summaries capture the reference summary's meaning, unlike surface-overlap metrics.

About these practice questions

This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.