NCA-GENL Trustworthy AI Practice Question
An AI team is preparing to release an LLM-powered legal research assistant. Before launch, they want to quantify how often the model produces confident but unsupported legal citations. Which evaluation approach most directly measures this failure mode?
⚠ Common exam trap
The trap here is substituting a fluency or sentiment proxy such as perplexity or user trust for a grounding check, when only source verification measures citation hallucination.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run a benchmark of legal queries and score each generated citation against an authoritative case-law database to compute a hallucination rate.
Directly testing whether generated citations exist and support their propositions requires checking them against an authoritative case-law source. Aggregating those checks yields a concrete hallucination rate that can be compared across versions and prompts. Response length, user surveys, and perplexity all measure adjacent qualities but cannot verify factual grounding, so they would not quantify the specific failure mode.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Measure the average response length in tokens across a sample of legal queries.
Why it's wrong here
Response length says nothing about whether cited cases exist or support the stated proposition. A model can produce long, citation-dense answers that are entirely fabricated, or short answers that are fully grounded. Token count is a stylistic metric, not a correctness metric. This approach would not distinguish a trustworthy citation from a hallucinated one, so it cannot quantify the failure mode the team is targeting.
- ✗
Compare the model's perplexity on a held-out set of legal documents to its perplexity on general web text.
Why it's wrong here
Perplexity measures how well a model predicts a text distribution and is not a reliable indicator of factual accuracy or citation validity. A model can have low perplexity on legal text while still fabricating case names that fit the stylistic pattern. Perplexity also says nothing about whether a cited authority supports a claim. This metric cannot isolate the hallucination behavior the team needs to measure.
- ✗
Survey the legal team to ask whether they generally trust the assistant's answers.
Why it's wrong here
Subjective trust surveys capture user sentiment, which can be influenced by writing style, speed, and confirmation bias. They do not produce a verifiable count of fabricated citations and may undercount errors that look plausible to a busy reviewer. A survey is useful for adoption research but not for quantifying a specific factual failure mode. The team needs an objective, per-citation check against authoritative sources, not an opinion poll.
- ✓
Run a benchmark of legal queries and score each generated citation against an authoritative case-law database to compute a hallucination rate.
Why this is correct
Grounding evaluation against an authoritative case-law database directly tests whether each cited case exists and matches the proposition it is attached to. Aggregating the results yields a hallucination rate, which is the quantity the team wants before launch. This approach targets the exact failure mode, confident but unsupported citations, and produces a metric that can be tracked across model versions and prompt changes.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.