CCAR-P Practice Question: Developer Productivity and Operational Enablement
A platform team maintains a shared prompt library used by 40 internal services. After a subtle wording change to a summarization prompt caused a 12% drop in a downstream classification F1 score, the team wants every prompt change to be reviewable, version-pinned, and automatically regression-tested before rollout. Which approach best satisfies these requirements?
⚠ Common exam trap
The trap here is assuming that ordinary unit tests with mocked model responses validate prompt quality, when only evaluation against real model outputs on a golden dataset can detect a semantic regression.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store prompts as versioned files in a Git repository, require pull-request review, and run a CI job that evaluates each changed prompt against a held-out golden dataset before merge.
Treating prompts as versioned code in Git couples three needed controls: peer review through pull requests, immutable references through commits or tags, and automated quality gates through a CI evaluation against a golden dataset. Because the evaluation runs before merge, a wording change that harms the downstream classification score is blocked rather than discovered in production, which protects all consuming services.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Move all prompts into a shared spreadsheet with a change log column, and instruct service owners to copy the latest text into their code before each release.
Why it's wrong here
A spreadsheet offers no enforced review, no immutable version identifier, and no automated evaluation, so a bad edit propagates silently. Manual copying into each service also guarantees drift between consumers and defeats version pinning. It cannot detect the 12% F1 regression before rollout, which is the core requirement here.
- ✗
Deploy every prompt change straight to production behind a feature flag, then compare aggregate F1 scores over the following week and revert manually if quality declines.
Why it's wrong here
This tests on live traffic, so users experience the degraded classification before anyone notices. A week-long observation window is far slower than a pre-merge evaluation and offers no peer review or immutable version history. Manual reverts also risk reintroducing the fault, and feature flags do not by themselves version prompt content.
- ✓
Store prompts as versioned files in a Git repository, require pull-request review, and run a CI job that evaluates each changed prompt against a held-out golden dataset before merge.
Why this is correct
Git gives immutable versions, diffable history, and mandatory peer review, while the CI evaluation job gates merge on measured quality against a golden dataset. This directly addresses reviewability, version pinning through commit SHAs or tags, and automated regression detection, so a wording change that degrades the classification F1 score is caught before it reaches the 40 consuming services.
- ✗
Keep prompts inline in each service's source code and rely on each team's existing unit tests, which mock the Claude response, to catch quality regressions.
Why it's wrong here
Inline prompts are duplicated across services and diverge over time, so there is no single reviewable artifact. Mocked responses test code paths, not model output quality, so a real 12% F1 drop would pass unnoticed. This provides neither centralized review nor genuine regression testing against model behavior.
About these practice questions
One of 262 original CCAR-P practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.