AIF-C01 Fundamentals of Generative AI Practice Question
A product team wants to compare two foundation models on their own customer-support transcripts before choosing one for a chatbot. They need a quantitative measure of how well each model's answers match reference answers. Which evaluation approach fits this need?
⚠ Common exam trap
A common mix-up: candidates confuse operational metrics such as latency and cost with quality metrics that actually measure answer correctness.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Automated evaluation using similarity metrics such as ROUGE or BERTScore between model outputs and reference answers
Automated text-similarity metrics score generated answers against reference answers and yield numbers that can be compared across models on the same transcripts. Human ratings are qualitative and costly, latency and cost measure operations rather than quality, and public benchmarks do not reflect the company's domain data. Automated evaluation on in-domain data is the fit.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Human evaluation where reviewers rate each answer on a five-point scale
Why it's wrong here
Human review captures nuance and is valuable for subjective quality, but it is slow, costly, and subject to rater disagreement. With a large transcript set it does not deliver the quick, repeatable quantitative comparison the team wants, and scaling it across two models multiplies the effort.
- ✓
Automated evaluation using similarity metrics such as ROUGE or BERTScore between model outputs and reference answers
Why this is correct
Automated metrics compare generated text against reference answers and produce numeric scores, enabling objective side-by-side comparison across models on the same dataset. This suits the requirement for a quantitative measure over the team's own transcripts, and it scales to many examples without manual grading.
- ✗
Reviewing each model's published model card and benchmark leaderboard rankings
Why it's wrong here
Model cards and public leaderboards describe general capabilities and training data, but they are not measured on the company's own transcripts. Domain-specific phrasing and intents differ from public benchmarks, so these sources cannot provide the tailored quantitative comparison the team requires.
- ✗
Measuring inference latency and cost per thousand tokens for each model
Why it's wrong here
Latency and cost are operational metrics, not quality measures. They say nothing about whether answers match reference responses, so a fast cheap model could still be inaccurate. These figures complement quality evaluation but cannot decide which model answers support questions better.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.