Courseiva

CCAR-P Governance, Safety, and Risk Management Practice Question

A media company uses Claude to moderate user-generated comments at high volume. The risk team wants a control that detects when the moderation model's behavior drifts, for example becoming unusually permissive or aggressive, before it affects the community at scale. Which control best fits this need?

⚠ Common exam trap

The trap here is equating operational metrics or reactive user feedback with behavioral drift detection, when only repeated comparison against a fixed labeled benchmark exposes a shift in model judgment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Maintain a labeled evaluation set of representative comments and run scheduled regression tests that compare current moderation decisions against expected outcomes, alerting on threshold breaches.

Detecting behavioral drift requires a stable reference: a labeled evaluation set that defines expected moderation outcomes. Running scheduled regression tests against it converts drift into a measurable deviation and supports alerting before community-wide impact. Token monitoring, user appeals, and verdict caching each serve different purposes and cannot reveal systematic changes in how the model judges content.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Provide an appeal button so users can report comments they believe were moderated incorrectly.

    Why it's wrong here

    Appeals surface individual errors but depend on users noticing and reporting them, which is slow and biased toward visible failures. A gradual, system-wide drift in moderation strictness may generate few appeals before causing broad harm. This is a reactive remediation channel, not a proactive detector of model behavior change.

  • ✓

    Maintain a labeled evaluation set of representative comments and run scheduled regression tests that compare current moderation decisions against expected outcomes, alerting on threshold breaches.

    Why this is correct

    A stable labeled evaluation set with scheduled regression runs turns drift into a measurable signal, comparing present behavior against known-good expectations. Alerts on threshold breaches catch permissiveness or aggression shifts before they spread across the community. The other options address cost, latency, or individual appeals rather than detecting systematic behavioral change over time.

  • ✗

    Monitor average tokens per moderation request to detect changes in comment length.

    Why it's wrong here

    Token counts reveal input size, not whether the model's moderation judgments have changed. A drift toward permissiveness or aggression produces no distinctive token signature, so this metric would stay flat while behavior shifted. It is a capacity-planning signal, not a behavioral quality control, and would miss the risk the team is trying to catch.

  • ✗

    Cache moderation verdicts for identical comments to reduce repeated inference costs.

    Why it's wrong here

    Caching reuses prior decisions for duplicate inputs, which improves efficiency but does not reveal whether the model's judgments are drifting. Worse, cached verdicts can mask behavioral change for repeated content while leaving novel comments unmeasured. Cost optimization is unrelated to the early-warning capability the risk team requested.

About these practice questions

This CCAR-P question is part of Courseiva's 262-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Anthropic exam blueprint

This CCAR-P practice question is part of Courseiva's free Anthropic certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CCAR-P exam.