CCAR-P Governance, Safety, and Risk Management Practice Question
A company is using Claude to process customer feedback. They want to ensure that if a customer mentions self-harm or illegal activities, the system immediately flags this for a human moderator. Which tool is best suited for this specific governance task?
⚠ Common exam trap
Candidates often suggest prompt engineering or 'system instructions' to handle safety. While helpful, these are insufficient for critical safety events; external moderation tools are required for reliable, auditable detection.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
An external Moderation API or a dedicated safety-tuned model layer.
Handling sensitive content like self-harm requires specialized safety tools that go beyond standard text classification. Anthropic and its partners provide moderation APIs and safety filters designed to detect these high-risk categories, allowing organizations to implement mandatory human intervention for critical safety events.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The 'Temperature' parameter, set to its lowest possible value.
Why it's wrong here
Temperature controls the randomness of the model's output and has no capability to detect or flag specific sensitive topics in the user's input. Using temperature for safety is a misunderstanding of how model parameters function in an inference pipeline.
- ✓
An external Moderation API or a dedicated safety-tuned model layer.
Why this is correct
Moderation APIs are specifically built to categorize text into safety buckets like 'self-harm', 'violence', or 'hate speech'. By routing customer feedback through a moderation layer before or alongside Claude, the system can trigger immediate alerts and human reviews for any dangerous content.
- ✗
A standard SQL database with a list of 'bad words' to block.
Why it's wrong here
Keyword blocking is often ineffective because it cannot understand context or intent. A user might use 'bad words' in a non-harmful way, or describe a dangerous situation without using any specific blacklisted terms, making this a weak and unreliable governance mechanism.
- ✗
Increasing the 'max_tokens' to allow the model to explain the risks.
Why it's wrong here
Increasing the token limit just allows for longer responses; it does not change the model's ability to flag content for external moderators. Relying on the model to 'explain risks' to a potentially vulnerable user is also less safe than immediate human intervention.
About these practice questions
This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Anthropic exam blueprint
This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.