Courseiva
Trustworthy AI →hardMultiple Select

NCA-GENL Trustworthy AI Practice Question

A healthcare organization is preparing to deploy an LLM-based clinical documentation assistant. The Trustworthy AI review board requires evidence that the model's outputs are safe and reliable before go-live. Which two practices should the team implement to provide this evidence? (Choose two.)

⚠ Common exam trap

The trap here is treating fine-tuning or a brief shadow deployment as sufficient proof of safety, when neither produces the structured, repeatable evidence a Trustworthy AI review requires.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Establish a continuous evaluation harness that scores model outputs against a held-out clinical benchmark and tracks metrics over time.

Adversarial red-teaming with clinicians and a continuous evaluation harness together provide both qualitative and quantitative evidence of safety and reliability. Red-teaming uncovers edge-case failures, while the harness tracks performance against a clinical benchmark over time. Combined, they give the review board documented, repeatable assurance that the assistant behaves safely across diverse patient scenarios.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Establish a continuous evaluation harness that scores model outputs against a held-out clinical benchmark and tracks metrics over time.

    Why this is correct

    A continuous evaluation harness provides quantitative, repeatable evidence of model quality on a representative clinical benchmark. Tracking metrics over time detects regressions and drift after updates, giving the review board ongoing assurance rather than a one-time snapshot. This is a core practice for demonstrating reliable performance in a regulated healthcare deployment.

  • ✗

    Increase the model's temperature setting to encourage more creative and varied clinical documentation.

    Why it's wrong here

    Raising temperature increases randomness, which is the opposite of what a clinical documentation assistant needs. Higher variability makes outputs less deterministic and harder to validate, and it raises hallucination risk in a domain where factual precision is critical. This would undermine rather than support the evidence of safety and reliability the review board requires.

  • ✗

    Fine-tune the model on a small curated dataset of ideal clinical notes and assume the fine-tuning eliminates all unsafe outputs.

    Why it's wrong here

    Fine-tuning on a small curated set can improve style and format but cannot guarantee elimination of unsafe outputs, especially for edge cases not represented in the data. Assuming it does removes the need for validation, which is precisely what the review board wants to see. This practice provides no measurable evidence of safety.

  • ✓

    Conduct adversarial red-teaming with clinicians to probe for unsafe or biased outputs across diverse patient scenarios.

    Why this is correct

    Red-teaming with domain experts surfaces failure modes that automated benchmarks miss, such as unsafe recommendations in rare clinical contexts or biased language across patient demographics. It produces documented evidence of the model's behavior under adversarial conditions, which directly supports the review board's requirement for demonstrated safety and reliability before deployment.

  • ✗

    Deploy the model in shadow mode for a single day and rely on informal developer feedback as the sole safety assessment.

    Why it's wrong here

    A one-day shadow deployment with informal feedback is not rigorous evidence. It lacks structured metrics, diverse scenario coverage, and expert review, so it cannot demonstrate safety or reliability to a review board. Shadow deployment can be a useful complement, but alone it is insufficient and the short duration limits the failure modes that would surface.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.