Courseiva

CCAR-P Governance, Safety, and Risk Management Practice Question

A security architect is performing red-teaming on a new Claude-powered application. They are specifically testing for 'jailbreaking' attempts where a user tries to bypass safety filters by using roleplay or adversarial framing. Which TWO strategies are most effective for mitigating this specific risk at the architectural level?

⚠ Common exam trap

Candidates often select client-side or prompt-only solutions, assuming standard instructions are bulletproof. They forget that jailbreaking specifically targets and bypasses text-based prompts, requiring multi-layered architectural safeguards like independent validators.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implementing a robust, immutable System Prompt

Mitigating adversarial attacks requires a multi-layered approach that combines model-native features with external validation. Using a strong system prompt sets clear boundaries that the model prioritizes, while implementing an independent moderation layer provides a final check on outputs. These strategies ensure that even if one layer is bypassed, the overall system remains resilient against malicious intent.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increasing the temperature parameter to 1.0

    Why it's wrong here

    Increasing temperature increases the randomness and creativity of the model's output, which can actually make its behavior less predictable and more susceptible to following adversarial prompts. High temperature settings generally degrade the reliability of safety guardrails and should be avoided in scenarios where strict adherence to safety is required.

  • ✓

    Implementing a robust, immutable System Prompt

    Why this is correct

    A well-defined system prompt acts as a foundational governance layer that defines the model's persona and safety constraints. By explicitly instructing the model to reject roleplay attempts that violate safety policies, architects can significantly harden the application against common jailbreaking techniques that rely on tricking the model into ignoring its rules.

  • ✗

    Reducing the max_tokens limit for all users

    Why it's wrong here

    Limiting the length of the response might truncate a harmful output, but it does not prevent the model from generating the harmful content in the first place. This is a superficial control that does not address the underlying vulnerability to adversarial framing or provide any real governance over model behavior.

  • ✓

    Utilizing a separate 'Safety' instance of Claude for output validation

    Why this is correct

    Deploying a secondary model to review the primary model's responses creates an independent validation loop. This 'checker' model can be specifically prompted to look for safety violations or policy bypasses, providing a robust second line of defense that is harder for a single adversarial prompt to compromise simultaneously.

  • ✗

    Switching from Claude 3.5 Sonnet to Claude 3 Haiku

    Why it's wrong here

    While Claude 3 Haiku is faster and more cost-effective, it generally has less reasoning capability than Sonnet, which may actually make it more prone to certain types of sophisticated adversarial manipulation. Choosing a smaller model is not a valid security strategy for mitigating jailbreaking risks in complex applications.

About these practice questions

One of 262 original CCAR-P practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.