AIF-C01 Applications of Foundation Models Practice Question
A company uses Amazon Bedrock to generate marketing copy. They want to measure the quality of generated text compared to reference text. Which metric is most appropriate?
⚠ Common exam trap
AWS often tests the distinction between classification/regression metrics and text generation metrics, leading candidates to mistakenly apply F1 score or accuracy to evaluate generated text quality instead of using BLEU or similar sequence-based metrics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
BLEU
BLEU (Bilingual Evaluation Understudy) is the most appropriate metric for evaluating the quality of generated text against reference text in tasks like machine translation and text generation. It measures n-gram precision between the generated and reference texts, making it ideal for assessing marketing copy generated by Amazon Bedrock.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
F1 score
Why it's wrong here
F1 score balances precision and recall for classification labels, so it cannot compare generated prose against reference text. It is tempting because F1 is a familiar evaluation metric, and it would be correct for judging a classifier's accuracy on labelled categories rather than text-generation quality.
- ✓
BLEU
Why this is correct
BLEU compares generated text against reference text using n-gram overlap, producing a precision-based score. It suits marketing copy evaluation where reference outputs exist. Perplexity, by contrast, measures language model likelihood without references, so it cannot assess similarity to target text.
- ✗
RMSE
Why it's wrong here
RMSE measures numeric prediction error, not text similarity, so it cannot score generated prose against references. It is tempting because it is a familiar regression metric, but the correct choice is a natural-language metric such as BERTScore or ROUGE.
- ✗
Accuracy
Why it's wrong here
Accuracy measures exact token-level matches, which fails for generative text evaluation because marketing copy can be semantically equivalent yet use different phrasing, so Bedrock’s output would be penalised for valid paraphrases. It is tempting because accuracy is a standard metric for classification tasks, such as verifying whether a model’s label matches a ground-truth category in a multiple-choice or binary outcome scenario.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.