Courseiva
easyMultiple Choice

AIF-C01 Practice Question: Evaluate the quality of a text generation model…

A company wants to evaluate the quality of a text generation model for a summarization task. They have reference summaries written by humans. Which automated metric compares the generated summary to the reference by measuring n-gram overlap?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ROUGE

ROUGE (Recall-Oriented Understudy for Gisting Evaluation) measures n-gram overlap between generated and reference summaries. BLEU is for translation, BERTScore uses embeddings, and perplexity measures language model confidence.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    BLEU

    Why it's wrong here

    BLEU measures precision of n-gram overlap, penalising length via brevity penalty, but it does not account for recall of the reference content, so it under-rewards summaries that capture all key points. It is designed for machine translation, where precision-oriented scoring of short outputs against references is the standard evaluation approach.

  • ✗

    Perplexity

    Why it's wrong here

    Perplexity scores how fluently a language model predicts a token sequence, requiring no reference summary at all, so it cannot compare output against human references. It is tempting because it is a standard intrinsic metric for language models, and would be correct when judging a model's fluency or comparing trained checkpoints.

  • ✓

    ROUGE

    Why this is correct

    ROUGE measures n-gram overlap between generated and reference summaries, directly satisfying the stem's requirement for an automated metric with human reference summaries. Recall-oriented variants such as ROUGE-N and ROUGE-L quantify how much reference content the generated summary captures, making it the standard choice for summarization evaluation.

  • ✗

    BERTScore

    Why it's wrong here

    BERTScore embeds tokens and computes cosine similarity between contextual vectors, so it never counts n-gram overlap; it captures semantic equivalence instead. It is tempting because it handles paraphrasing and synonymy that BLEU penalises, making it the right pick when generated text differs in wording but matches reference meaning.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.