AIF-C01 Applications of Foundation Models Practice Question
A startup prototypes a support assistant using Amazon Bedrock and needs to compare how three different foundation models handle the same set of 200 support tickets. They want objective quality scores, including accuracy against reference answers and robustness, before committing to one model. Which approach should they use?
⚠ Common exam trap
The trap here is substituting a safety-control signal or a performance metric for genuine answer-quality evaluation against reference responses.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create an Amazon Bedrock model evaluation job with the tickets and reference answers.
The startup needs standardized, repeatable quality metrics across multiple models. Amazon Bedrock model evaluation jobs accept a dataset with prompts and reference responses and compute metrics such as accuracy, robustness, and toxicity for each selected model, making results directly comparable. Guardrail block rates, informal human ratings, and latency measurements capture different properties and cannot substitute for objective quality scoring.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create an Amazon Bedrock model evaluation job with the tickets and reference answers.
Why this is correct
Amazon Bedrock model evaluation jobs run a dataset of prompts and reference responses against selected models and produce metrics such as accuracy, robustness, and toxicity. Running the 200 tickets through one job gives comparable, objective scores across the three models, which is exactly what the startup needs to choose a model before committing.
- ✗
Increase Provisioned Throughput for each model and measure response latency.
Why it's wrong here
Provisioned Throughput and latency measurement describe performance and capacity characteristics, not answer correctness. A fast model may still be inaccurate. This approach ignores the requirement to score accuracy against reference answers and robustness, so it would not inform the model choice.
- ✗
Enable Guardrails for Amazon Bedrock on each model and compare block rates.
Why it's wrong here
Guardrails reports how often content is filtered, which reflects policy violations rather than answer quality. A model could block nothing yet answer inaccurately. Block-rate comparison does not provide accuracy or robustness scores against reference answers, so it cannot drive the model selection decision described.
- ✗
Send the tickets through each model and have engineers rate the replies informally.
Why it's wrong here
Informal human ratings are subjective, inconsistent across reviewers, and difficult to reproduce. They do not yield standardized accuracy or robustness metrics, so comparisons between models would be unreliable. The scenario explicitly asks for objective scores, which this approach cannot provide.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.