AIF-C01 Fundamentals of Generative AI Practice Question
A support organization notices its generative AI assistant sometimes invents policy details that do not exist. Leadership wants a measurable way to track whether this behavior improves over time. Which action should the team take?
⚠ Common exam trap
The trap here is assuming that newer, larger models or more output space automatically fix hallucination, when without a measurement framework no one can tell whether the problem improved.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Build a labeled evaluation set of questions with verified answers and score responses for factual consistency on a recurring basis.
Turning an observed quality problem into a tracked metric requires a fixed evaluation set with verified answers and a scoring procedure that can be rerun after each change. This makes comparisons across prompts, retrieval configurations, and models meaningful, and it establishes a regression baseline so the organization can prove whether fabrication is actually declining.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to a model with a larger parameter count and assume the fabrication problem disappears.
Why it's wrong here
Larger models can still hallucinate, especially when asked about organization-specific policies absent from training data, and the change provides no way to verify whether the problem improved. Without a measurement framework, the team cannot demonstrate progress or detect regressions after the swap.
- ✓
Build a labeled evaluation set of questions with verified answers and score responses for factual consistency on a recurring basis.
Why this is correct
A curated evaluation set with known-correct answers turns an anecdotal complaint into a repeatable metric, allowing the team to compare prompts, retrieval strategies, or models over time. This directly addresses the request for a measurable way to track improvement in factual accuracy and supports regression testing as the system evolves.
- ✗
Increase the model's maximum output token limit so it has more room to explain itself.
Why it's wrong here
Allowing longer outputs gives the model more space to elaborate, which frequently increases the number of unsupported claims rather than reducing them. Output length is unrelated to whether statements are factually grounded, so this change would not provide a measurement mechanism and may worsen the observed behavior.
- ✗
Rely on customer complaints as the primary signal for whether fabricated policy details are decreasing.
Why it's wrong here
Complaint volume is lagging, noisy, and biased toward the small fraction of errors users notice and bother to report. It cannot distinguish a real reduction in fabrication from changes in traffic or reporting behavior, so it fails as a measurement instrument for tracking accuracy improvements over time.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.