mediumMultiple Choice
AIF-C01 Practice Question: A team is evaluating two different foundation…
A team is evaluating two different foundation models for a sentiment analysis task. They have a labeled test dataset. Which evaluation approach should they use to compare the models' performance on this task?
⚠ Common exam trap
AWS often tests the distinction between metrics for generative tasks (ROUGE, BLEU) versus classification tasks (accuracy, precision, recall, F1), leading candidates to mistakenly apply text-generation metrics to a classification problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compute accuracy, precision, recall, and F1-score on the labeled test set
Accuracy, precision, recall, and F1-score are standard classification metrics that directly measure how well a model's predicted sentiment labels match the ground truth labels in a labeled test dataset. These metrics provide a quantitative, reproducible comparison of model performance on a supervised sentiment analysis task, unlike text-generation metrics or subjective human evaluation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use ROUGE scores to compare model outputs
Why it's wrong here
ROUGE compares generated summaries against reference text via n-gram and sequence overlap, so it does not evaluate discrete sentiment labels. It is tempting because it is a familiar NLP metric, but a labelled test set for classification calls for accuracy, precision, recall, or F1 instead.
- ✗
Run a human evaluation study with 100 judges
Why it's wrong here
Human evaluation lacks the mathematical objectivity required to calculate quantitative metrics like F1-score or accuracy against a pre-existing labeled test dataset. This method relies on subjective qualitative judgement rather than statistical comparison against ground truth labels. You would utilise human judges when assessing generative qualities, such as hallucination rates or stylistic nuance, where no definitive gold standard exists to automate the scoring process.
- ✗
Use BLEU scores to measure n-gram overlap with reference labels
Why it's wrong here
BLEU is for translation, not sentiment analysis.
- ✓
Compute accuracy, precision, recall, and F1-score on the labeled test set
Why this is correct
A labelled test set supports supervised classification metrics, so computing accuracy, precision, recall, and F1-score quantifies each model's predictions against ground-truth sentiment labels. This gives an objective, comparable measurement of both models on the same task.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.