Courseiva
Trustworthy AI →mediumMultiple Choice

NCA-GENL Trustworthy AI Practice Question

An AI team is preparing to release an LLM-powered legal research assistant. Before launch, they want to quantify how often the model produces confident but unsupported legal citations. Which evaluation approach most directly measures this failure mode?

⚠ Common exam trap

The trap here is substituting a fluency or sentiment proxy such as perplexity or user trust for a grounding check, when only source verification measures citation hallucination.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run a benchmark of legal queries and score each generated citation against an authoritative case-law database to compute a hallucination rate.

Directly testing whether generated citations exist and support their propositions requires checking them against an authoritative case-law source. Aggregating those checks yields a concrete hallucination rate that can be compared across versions and prompts. Response length, user surveys, and perplexity all measure adjacent qualities but cannot verify factual grounding, so they would not quantify the specific failure mode.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Measure the average response length in tokens across a sample of legal queries.

    Why it's wrong here

    Response length says nothing about whether cited cases exist or support the stated proposition. A model can produce long, citation-dense answers that are entirely fabricated, or short answers that are fully grounded. Token count is a stylistic metric, not a correctness metric. This approach would not distinguish a trustworthy citation from a hallucinated one, so it cannot quantify the failure mode the team is targeting.

  • ✗

    Compare the model's perplexity on a held-out set of legal documents to its perplexity on general web text.

    Why it's wrong here

    Perplexity measures how well a model predicts a text distribution and is not a reliable indicator of factual accuracy or citation validity. A model can have low perplexity on legal text while still fabricating case names that fit the stylistic pattern. Perplexity also says nothing about whether a cited authority supports a claim. This metric cannot isolate the hallucination behavior the team needs to measure.

  • ✗

    Survey the legal team to ask whether they generally trust the assistant's answers.

    Why it's wrong here

    Subjective trust surveys capture user sentiment, which can be influenced by writing style, speed, and confirmation bias. They do not produce a verifiable count of fabricated citations and may undercount errors that look plausible to a busy reviewer. A survey is useful for adoption research but not for quantifying a specific factual failure mode. The team needs an objective, per-citation check against authoritative sources, not an opinion poll.

  • ✓

    Run a benchmark of legal queries and score each generated citation against an authoritative case-law database to compute a hallucination rate.

    Why this is correct

    Grounding evaluation against an authoritative case-law database directly tests whether each cited case exists and matches the proposition it is attached to. Aggregating the results yields a hallucination rate, which is the quantity the team wants before launch. This approach targets the exact failure mode, confident but unsupported citations, and produces a metric that can be tracked across model versions and prompt changes.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.