CCAR-P Practice Question: Developer Productivity and Operational Enablement
A team uses a CI/CD pipeline to deploy LLM applications. They want to ensure prompt changes do not degrade model performance. Which strategy best integrates evaluation into the development workflow?
⚠ Common exam trap
Candidates often suggest periodic manual audits or post-deployment monitoring, overlooking the critical requirement to integrate automated evaluation gates directly into the CI/CD pipeline for immediate feedback.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run automated evaluation scripts against a golden dataset during the CI process.
Integrating automated evaluations (Evals) into the CI/CD pipeline ensures that every code or prompt change is validated against a golden dataset. This automated gate prevents regressions from reaching production. It empowers developers to iterate quickly while maintaining a high bar for reliability, which is crucial for operational enablement in LLM-driven environments where non-deterministic model behavior can introduce subtle, hard-to-detect bugs that impact user experience.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Perform manual prompt testing by the QA team before every release.
Why it's wrong here
Manual QA does not scale in an agile development environment and introduces significant bottlenecks. It is prone to human error and lacks the statistical rigor needed for testing generative models. Automating evaluations is necessary for achieving the continuous deployment speed expected in modern software development cycles.
- ✓
Run automated evaluation scripts against a golden dataset during the CI process.
Why this is correct
Automated evaluation against a golden dataset provides objective, repeatable metrics to validate prompt quality before deployment. This allows developers to catch regressions early in the SDLC. By integrating this into CI, teams maintain high deployment velocity without sacrificing the quality or safety of the LLM application outputs.
- ✗
Rely on user feedback loops in production to identify and fix issues.
Why it's wrong here
Relying on user feedback means customers are the ones detecting bugs, which negatively impacts user experience and brand reputation. This reactive approach is inefficient and dangerous for mission-critical applications. Proactive, automated testing is essential to ensure that prompt modifications are safe and effective before they ever reach the user.
- ✗
Only evaluate the application after it has been deployed to the production environment.
Why it's wrong here
Post-deployment evaluation is too late to prevent issues from affecting users. It creates a higher risk environment where rollbacks are expensive and complex. Evaluation must be 'shifted left' into the CI/CD pipeline to ensure that problematic changes are intercepted before they become a production incident.
About these practice questions
Courseiva writes every CCAR-P question from scratch — 262 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.