CCAR-P Practice Question: Developer Productivity and Operational Enablement
A developer is building a tool that lets engineers query an internal knowledge base through Claude. During testing, Claude sometimes invents plausible but nonexistent document titles when the retrieved context is thin. The team wants a systematic way to detect and reduce these hallucinations before the tool reaches general availability. Which approach is most appropriate?
⚠ Common exam trap
The trap here is reaching for a sampling parameter or a stern instruction to cure fabrication, when hallucination of this kind stems from insufficient or poorly ranked grounding context.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Build an evaluation set of questions with known ground-truth answers, score responses for faithfulness to the retrieved context, and iterate on retrieval and prompting until scores meet a defined threshold.
Hallucinated titles are a grounding failure, so the fix is to measure faithfulness against retrieved context and improve retrieval and prompting where scores fall short. A labeled evaluation set with ground-truth answers makes the problem quantifiable, and a release threshold converts that measurement into a defensible go or no-go decision.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add an instruction telling the model to be accurate and never hallucinate, then rely on that instruction in production.
Why it's wrong here
Generic accuracy instructions are weak controls and provide no measurement, so the team cannot tell whether fabrication decreased. The underlying cause, insufficient grounding context, remains unaddressed. Without an evaluation set or scoring, regressions go undetected, making this unsuitable as a systematic pre-release approach.
- ✗
Raise the temperature setting so responses become more varied and the model is less likely to repeat a fabricated title.
Why it's wrong here
Higher temperature increases randomness and generally amplifies fabrication rather than reducing it. It also makes outputs less reproducible, which undermines systematic detection. This option treats a grounding problem as a sampling problem, so it fails to address why the model invents titles when retrieved context is thin.
- ✓
Build an evaluation set of questions with known ground-truth answers, score responses for faithfulness to the retrieved context, and iterate on retrieval and prompting until scores meet a defined threshold.
Why this is correct
A labeled evaluation set with faithfulness scoring turns a vague symptom into a measurable signal. Iterating on retrieval quality and prompting against that metric addresses the root cause, since thin context drives fabrication. A defined threshold provides an objective release gate, so the team can demonstrate improvement rather than relying on anecdotal spot checks.
- ✗
Switch the knowledge base queries to return a fixed number of documents regardless of relevance, so the model always has something to cite.
Why it's wrong here
Forcing a fixed document count injects irrelevant passages that the model may still misattribute, and it hides genuine retrieval gaps. It does not measure hallucination or improve grounding quality. The approach also wastes context window space, and it offers no way to verify improvement before general availability.
About these practice questions
Courseiva writes every CCAR-P question from scratch — 262 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.