Courseiva

CCAR-P Governance, Safety, and Risk Management Practice Question

An insurer uses a Claude-based agent that can call internal tools to look up policy details. During review, the safety team finds that a document uploaded by a claimant contains text instructing the agent to email the full policy database to an external address. The agent has an email tool available. Which control most directly prevents this class of failure?

⚠ Common exam trap

The trap here is treating prompt injection as a text-filtering problem, when the decisive control is limiting what the agent is authorized to do rather than what it is willing to read.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enforce tool-level authorization so the agent's identity lacks permission to send external email, and constrain tool arguments to validated allowlists.

When untrusted content can reach a model that holds real capabilities, the reliable fix is to remove the dangerous capability from the agent's identity and constrain the arguments of the tools it keeps. Injected text can persuade a model but cannot grant permissions the credential layer denies, so an agent without external-email rights cannot exfiltrate regardless of what the document says. Detection and prompting remain useful layers but are not deterministic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Scan uploaded documents with a classifier that flags imperative language and quarantine any document that scores above threshold.

    Why it's wrong here

    Detection helps surface obvious attacks and generates useful telemetry, but classifiers produce false negatives, and an attacker who knows the detector exists can rephrase to evade it. Quarantining also disrupts legitimate claims containing ordinary imperative text. Detection reduces the attack surface but cannot guarantee that a dangerous tool call never fires.

  • ✗

    Add a system prompt instruction telling the model to ignore any instructions found inside uploaded documents.

    Why it's wrong here

    Instructions to disregard embedded commands are a useful defense-in-depth layer, but they are probabilistic. A sufficiently crafted document can still influence a model that has already been told to be helpful with the document's contents. Because the consequence here is exfiltration of an entire database, a control that can be argued around in natural language is too weak to be the primary safeguard.

  • ✗

    Require human approval for every tool invocation the agent proposes, regardless of which tool or argument is involved.

    Why it's wrong here

    Universal approval eliminates autonomous action but destroys the value of the agent and creates review fatigue, which itself causes rubber-stamping. It is also disproportionate: most invocations are routine lookups with no exfiltration potential. A targeted authorization boundary on the dangerous capability is both stronger and more sustainable than blanket human gating.

  • ✓

    Enforce tool-level authorization so the agent's identity lacks permission to send external email, and constrain tool arguments to validated allowlists.

    Why this is correct

    Removing the agent's ability to send external mail makes the injected instruction inert no matter how persuasive the document is, because the capability simply does not exist at the credential layer. Argument allowlisting further blocks misuse of tools the agent does retain, such as restricting a lookup to the current claimant's policy. This is a deterministic boundary rather than a behavioral request.

About these practice questions

This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.