Courseiva

CCAR-P Governance, Safety, and Risk Management Practice Question

A financial services firm runs a Claude-powered agent that can call internal tools to move funds between accounts. Risk management wants a control that prevents the agent from executing a transfer above a threshold without human sign-off, and that remains effective even if the model is manipulated through injected content in a retrieved document. Which control best meets this requirement?

⚠ Common exam trap

The trap here is believing that a system prompt instruction or a second reviewing model constitutes an enforcement control, when only the component that executes the action can reliably refuse it.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enforce the threshold in the tool-execution layer so that any transfer exceeding the limit is rejected unless a human approval token is presented.

When an agent can take consequential action, the authorization boundary must live outside the model, in the component that executes the action. Enforcing the threshold and requiring a human approval token at the tool layer means a manipulated model cannot exceed its authority, because the code performing the transfer refuses. Prompt rules, reviewer models, and sampling adjustments modify model behavior but cannot guarantee that a hijacked agent will decline a prohibited action.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Lower the agent's temperature and restrict its tool list to a single transfer function with a fixed daily cap.

    Why it's wrong here

    A daily cap limits aggregate exposure but does not prevent a single unauthorized transfer above the per-transaction threshold, which is the stated requirement. Lowering temperature reduces randomness, not adversarial influence, and narrowing the tool list does not add an approval gate. The control must reject the specific over-threshold action absent human sign-off, and none of these measures does that.

  • ✓

    Enforce the threshold in the tool-execution layer so that any transfer exceeding the limit is rejected unless a human approval token is presented.

    Why this is correct

    Placing the limit in the layer that actually performs the transfer makes the control independent of model behavior, so a manipulated agent still cannot move funds above the threshold. Requiring a human approval token binds the exception to an accountable person. This is defense in depth: the model may propose, but the execution layer disposes, which is the only design that survives prompt injection.

  • ✗

    Add a system prompt rule stating that transfers above the threshold must be escalated to a human operator.

    Why it's wrong here

    Instructions in a system prompt shape model behavior but are not an enforcement boundary; injected content in a retrieved document can persuade the model to disregard them, which is precisely the threat described. The control must hold even when the model is compromised, and a prompt-level rule cannot make that guarantee. It documents intent rather than preventing the action.

  • ✗

    Run a second Claude instance as a reviewer that inspects each proposed transfer and vetoes suspicious ones.

    Why it's wrong here

    A model-based reviewer adds a probabilistic check that can itself be manipulated, especially when the injected content is part of the reviewed context. It also adds latency and cost without creating a hard boundary. Because both the actor and the reviewer reason over the same tainted inputs, a sufficiently crafted injection can pass both, leaving the funds transfer unprotected.

About these practice questions

This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.