Courseiva

CCAR-P Governance, Safety, and Risk Management Practice Question

A company is using Claude to process customer feedback. They want to ensure that if a customer mentions self-harm or illegal activities, the system immediately flags this for a human moderator. Which tool is best suited for this specific governance task?

⚠ Common exam trap

Candidates often suggest prompt engineering or 'system instructions' to handle safety. While helpful, these are insufficient for critical safety events; external moderation tools are required for reliable, auditable detection.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

An external Moderation API or a dedicated safety-tuned model layer.

Handling sensitive content like self-harm requires specialized safety tools that go beyond standard text classification. Anthropic and its partners provide moderation APIs and safety filters designed to detect these high-risk categories, allowing organizations to implement mandatory human intervention for critical safety events.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The 'Temperature' parameter, set to its lowest possible value.

    Why it's wrong here

    Temperature controls the randomness of the model's output and has no capability to detect or flag specific sensitive topics in the user's input. Using temperature for safety is a misunderstanding of how model parameters function in an inference pipeline.

  • ✓

    An external Moderation API or a dedicated safety-tuned model layer.

    Why this is correct

    Moderation APIs are specifically built to categorize text into safety buckets like 'self-harm', 'violence', or 'hate speech'. By routing customer feedback through a moderation layer before or alongside Claude, the system can trigger immediate alerts and human reviews for any dangerous content.

  • ✗

    A standard SQL database with a list of 'bad words' to block.

    Why it's wrong here

    Keyword blocking is often ineffective because it cannot understand context or intent. A user might use 'bad words' in a non-harmful way, or describe a dangerous situation without using any specific blacklisted terms, making this a weak and unreliable governance mechanism.

  • ✗

    Increasing the 'max_tokens' to allow the model to explain the risks.

    Why it's wrong here

    Increasing the token limit just allows for longer responses; it does not change the model's ability to flag content for external moderators. Relying on the model to 'explain risks' to a potentially vulnerable user is also less safe than immediate human intervention.

About these practice questions

This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.