Courseiva
mediumMultiple Choice

AIF-C01 Practice Question: A team is evaluating two different foundation…

A team is evaluating two different foundation models for a sentiment analysis task. They have a labeled test dataset. Which evaluation approach should they use to compare the models' performance on this task?

⚠ Common exam trap

AWS often tests the distinction between metrics for generative tasks (ROUGE, BLEU) versus classification tasks (accuracy, precision, recall, F1), leading candidates to mistakenly apply text-generation metrics to a classification problem.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Compute accuracy, precision, recall, and F1-score on the labeled test set

Accuracy, precision, recall, and F1-score are standard classification metrics that directly measure how well a model's predicted sentiment labels match the ground truth labels in a labeled test dataset. These metrics provide a quantitative, reproducible comparison of model performance on a supervised sentiment analysis task, unlike text-generation metrics or subjective human evaluation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use ROUGE scores to compare model outputs

    Why it's wrong here

    ROUGE compares generated summaries against reference text via n-gram and sequence overlap, so it does not evaluate discrete sentiment labels. It is tempting because it is a familiar NLP metric, but a labelled test set for classification calls for accuracy, precision, recall, or F1 instead.

  • ✗

    Run a human evaluation study with 100 judges

    Why it's wrong here

    Human evaluation lacks the mathematical objectivity required to calculate quantitative metrics like F1-score or accuracy against a pre-existing labeled test dataset. This method relies on subjective qualitative judgement rather than statistical comparison against ground truth labels. You would utilise human judges when assessing generative qualities, such as hallucination rates or stylistic nuance, where no definitive gold standard exists to automate the scoring process.

  • ✗

    Use BLEU scores to measure n-gram overlap with reference labels

    Why it's wrong here

    BLEU is for translation, not sentiment analysis.

  • ✓

    Compute accuracy, precision, recall, and F1-score on the labeled test set

    Why this is correct

    A labelled test set supports supervised classification metrics, so computing accuracy, precision, recall, and F1-score quantifies each model's predictions against ground-truth sentiment labels. This gives an objective, comparable measurement of both models on the same task.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.