AIF-C01 Applications of Foundation Models Practice Question
A retail company wants to compare the output quality of several foundation models available in Amazon Bedrock for a product description generation task. They need a repeatable, automated way to score responses against reference descriptions. Which AWS capability should they use?
⚠ Common exam trap
Many exam-takers confuse operational monitoring or capacity features with model quality evaluation, which requires scoring outputs against references.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon Bedrock model evaluation jobs with automatic metrics such as BERTScore and ROUGE.
Amazon Bedrock model evaluation jobs run automated scoring against reference datasets using metrics such as BERTScore and ROUGE, enabling consistent comparison of multiple foundation models. This directly supports selecting the best model for product description generation with repeatable, quantitative results rather than subjective manual review.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Amazon Bedrock model evaluation jobs with automatic metrics such as BERTScore and ROUGE.
Why this is correct
Amazon Bedrock supports model evaluation jobs that can automatically score model outputs against reference data using metrics like BERTScore and ROUGE. This provides a repeatable, automated comparison across multiple foundation models, which matches the requirement. It removes manual scoring and produces consistent results that can guide model selection for the product description task.
- ✗
AWS Trusted Advisor to assess foundation model performance.
Why it's wrong here
AWS Trusted Advisor provides recommendations on cost optimization, security, fault tolerance, and service limits; it does not evaluate generative model output quality. It has no capability to score text against references or compare foundation models. Using it for this purpose would not produce the automated quality metrics the company needs.
- ✗
Amazon Bedrock provisioned throughput to benchmark model latency.
Why it's wrong here
Provisioned throughput reserves model capacity and affects latency and cost, but it does not measure output quality against reference descriptions. The scenario is about scoring generated text quality across models, not about throughput or latency benchmarking. Provisioned throughput is an operational purchasing option, not an evaluation capability.
- ✗
Amazon CloudWatch Logs to capture model responses and manually review them.
Why it's wrong here
CloudWatch Logs can store invocation logs, but it does not compute quality scores or compare models automatically. Manual review is not repeatable or scalable for evaluating several models against reference descriptions. The scenario explicitly asks for an automated scoring method, so logging alone does not satisfy the requirement.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.