Courseiva
hardMultiple Choice

AIF-C01 Practice Question: A data scientist is evaluating two different…

A data scientist is evaluating two different foundation models for a summarization task. They want to compare the quality of summaries generated by each model against a set of human-written reference summaries. Which set of metrics is most appropriate for this automated evaluation?

⚠ Common exam trap

AWS often tests the distinction between evaluation metrics for classification (accuracy, F1) versus generation (ROUGE, BLEU, BERTScore), and candidates mistakenly apply classification metrics to summarization tasks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ROUGE, BLEU, and BERTScore

ROUGE, BLEU, and BERTScore are specifically designed for evaluating text generation quality against reference summaries. ROUGE measures n-gram overlap (recall-oriented), BLEU measures precision of n-gram matches, and BERTScore uses contextual embeddings from BERT to capture semantic similarity, making them ideal for summarization tasks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Accuracy, precision, recall, and F1-score

    Why it's wrong here

    Accuracy, precision, recall, and F1-score assume discrete class labels, whereas summarisation produces free text compared against references. They are tempting because they are the standard classification metrics, and they would be correct for a labelling or intent-detection task, but they cannot score overlapping n-grams or semantic similarity.

  • ✓

    ROUGE, BLEU, and BERTScore

    Why this is correct

    ROUGE, BLEU and BERTScore compare generated text against reference summaries using n-gram overlap and semantic similarity. This directly satisfies the stem's constraint of automated evaluation against human-written references, unlike classification metrics such as accuracy or F1.

  • ✗

    Latency and throughput

    Why it's wrong here

    Latency and throughput measure operational serving performance, not summary fidelity against human references. They are tempting because they matter when selecting a model for production deployment under throughput constraints, but they cannot detect whether generated summaries preserve the reference content, so ROUGE and BERTScore are needed instead.

  • ✗

    Mean squared error (MSE) and R-squared

    Why it's wrong here

    MSE and R-squared quantify error between continuous numeric predictions and observed values, which summarisation does not produce. They are tempting because they are the default regression metrics, and they would be correct for forecasting or price prediction, but text overlap and semantic similarity require ROUGE and BERTScore instead.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.