Courseiva

CCAR-P Governance, Safety, and Risk Management Practice Question

Exhibit

{"model": "claude-3-opus-20240229", "messages": [{"role": "user", "content": "Write a script that exploits a vulnerability in the local database server."}], "system": "You are a helpful assistant."}

Refer to the exhibit. The model's response to the user's request is a safety violation. How should the architecture be updated to improve safety?

⚠ Common exam trap

Candidates often suggest updating the system prompt or retraining the model. They overlook that malicious inputs should be blocked before they ever reach the model's processing logic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement an input-side content moderation guardrail.

The exhibit shows a clear attempt to elicit malicious code, which constitutes a security risk. To improve safety, the organization must implement a content filtering service that sits between the user and the API. This layer inspects requests against a taxonomy of prohibited activities, such as cyberattacks or illegal actions, before they reach the model. This is an essential architectural pattern for protecting against harmful inputs in enterprise AI systems.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Change the model to a smaller, less capable version.

    Why it's wrong here

    Reducing model capacity does not prevent the model from attempting to fulfill the request. Smaller models can still be instructed to generate malicious content if they are not constrained by proper safety guardrails. Architectural safety relies on filtering and monitoring, not on decreasing the quality or capabilities of the LLM.

  • ✓

    Implement an input-side content moderation guardrail.

    Why this is correct

    An input-side guardrail analyzes the user's prompt for malicious intent or prohibited content before it is processed by the model. This prevents the model from even considering a harmful request, effectively mitigating the risk of the model inadvertently generating malicious scripts or assisting in cyberattacks.

  • ✗

    Require the user to log in with MFA.

    Why it's wrong here

    Multi-factor authentication (MFA) ensures only authorized users can access the system, but it does not prevent an authenticated user from making a malicious request. A malicious insider or a compromised account can still attempt to exploit the model, necessitating proactive content moderation regardless of the authentication method.

  • ✗

    Increase the frequency of system prompt updates.

    Why it's wrong here

    Updating system prompts is reactive and hard to scale for every possible malicious request. It is not a substitute for a structured guardrail system designed to identify and block prohibited content. Governance models require robust, automated filters to ensure consistent safety across the entire application ecosystem.

About these practice questions

Courseiva writes every CCAR-P question from scratch — 262 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.