CCAR-P Practice Question: Developer Productivity and Operational Enablement
A team maintains a Claude-powered code review bot. Reviewers complain that the bot sometimes approves pull requests that clearly violate the team's security policy. The team wants to make policy violations detectable and reproducible in CI without relying on manual spot checks. What is the most effective approach?
⚠ Common exam trap
The trap here is treating a prompt tweak or a bigger model as a fix, when without labeled tests the team cannot prove the change works or catch future regressions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Build a regression suite of labeled pull requests with known policy outcomes, run the bot against it on every prompt or model change, and fail CI when accuracy drops below a threshold.
The only approach that makes violations detectable and reproducible is a labeled regression suite wired into CI with an explicit accuracy threshold. It converts anecdotal complaints into a metric, catches regressions on prompt or model changes, and gates deployment. Prompt tweaks, manual reporting, and model size upgrades all lack the measurement needed to prove the bot is improving.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Ask reviewers to report every false approval in a shared spreadsheet and review the log monthly.
Why it's wrong here
Manual reporting is slow, incomplete, and not reproducible in CI. Reviewers will miss cases, and monthly review means regressions persist for weeks. It creates no automated gate to block a bad prompt or model change. The team explicitly wants to move away from manual spot checks, so this approach contradicts the stated requirement.
- ✗
Switch to a larger model and rely on its improved reasoning to catch policy violations without additional testing.
Why it's wrong here
A larger model may improve reasoning, but without a labeled suite the team cannot verify it reduces missed violations, and it may change behaviour on other cases. Model upgrades can introduce new failure modes, so relying on size alone is not reproducible. This also increases cost and latency without giving CI a measurable quality gate.
- ✗
Raise the model temperature to zero and add 'be strict about security' to the system prompt, then redeploy.
Why it's wrong here
Lowering temperature and adding an instruction may reduce variance, but without a labeled test set the team cannot prove the change fixes the missed violations or that it does not introduce false positives. It is a guess rather than a measurable improvement. This approach also cannot detect future regressions when the model or prompt changes again.
- ✓
Build a regression suite of labeled pull requests with known policy outcomes, run the bot against it on every prompt or model change, and fail CI when accuracy drops below a threshold.
Why this is correct
A labeled regression suite turns vague quality complaints into a measurable signal. Running it on every prompt or model change catches regressions before they reach reviewers, and a CI threshold makes the quality bar explicit. This makes violations reproducible and detectable automatically, which is exactly what the team needs to trust the bot over time.
About these practice questions
This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.