hardMultiple Choice
AIF-C01 Practice Question: Using Amazon Bedrock to build a sentiment…
A company is using Amazon Bedrock to build a sentiment analysis application for customer reviews. They need to evaluate the model's performance against a labeled test dataset. They want to use a metric that compares the model's predicted sentiment (positive, negative, neutral) to the ground truth labels. Which metric is MOST appropriate?
⚠ Common exam trap
AWS often tests the misconception that advanced NLP metrics like BLEU or ROUGE are suitable for classification tasks, when in fact they are designed for generation tasks, leading candidates to overcomplicate the answer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Accuracy
Accuracy is the most appropriate metric for this multi-class classification task because it directly measures the proportion of correctly predicted sentiments (positive, negative, neutral) out of total predictions, given a labeled test dataset. It is simple, interpretable, and well-suited for evaluating categorical outputs where class distribution is balanced or the cost of misclassification is uniform.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
ROUGE-1
Why it's wrong here
ROUGE-1 measures unigram overlap between generated and reference summaries, which is meaningless for comparing categorical sentiment labels. It is intended for summarisation evaluation. A classification metric such as accuracy or F1 against the labelled test set is the appropriate measure.
- ✗
BERTScore
Why it's wrong here
BERTScore embeds tokens and measures semantic similarity between generated and reference sentences, so it does not evaluate a three-way label prediction. It suits text generation quality assessment. Comparing predicted sentiment labels with ground truth needs accuracy, precision, recall or F1.
- ✗
BLEU
Why it's wrong here
BLEU scores n-gram overlap between generated and reference text, so it cannot compare discrete class labels such as positive, negative and neutral. It is designed for machine translation and summarisation evaluation. Classification metrics like accuracy, precision or F1 against the labelled dataset are required here.
- ✓
Accuracy
Why this is correct
Accuracy directly measures the proportion of predicted sentiment labels matching ground truth across the three classes, giving a single comparable figure for the labelled test set. It satisfies the requirement to compare predictions against ground truth without weighting any class.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.