NCP-GENL Safety, Ethics, and Compliance Practice Question
A healthcare company deploys an LLM-powered clinical documentation assistant using NVIDIA NIM microservices on-premises. During a compliance review, auditors discover that the model occasionally generates patient names and medical record numbers in its output even though these were not present in the input prompt. The team needs to implement a runtime guardrail that detects and blocks such unintended PII leakage without retraining the model. Which NVIDIA component should they configure to add this output-side detection?
⚠ Common exam trap
Test-takers frequently confuse inference optimization or serving tools with runtime content moderation, assuming that any NVIDIA component in the pipeline can enforce PII policies.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NVIDIA NeMo Guardrails with an output rail that invokes a PII detection model
Output rails in NeMo Guardrails are designed to intercept the model's generated text and apply safety checks, such as PII detection, before the response is returned. This is the correct mechanism because it operates at runtime without modifying the underlying model. Triton, TensorRT-LLM, and Riva serve different purposes—serving, optimization, and speech—and lack content-filtering capabilities for PII.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
NVIDIA NeMo Guardrails with an output rail that invokes a PII detection model
Why this is correct
NeMo Guardrails supports output rails that run after the LLM generates a response, allowing a PII detection model to scan and block or redact sensitive data before it reaches the user. This directly addresses unintended PII leakage without retraining, as the guardrail operates at inference time and can be configured with custom detection logic.
- ✗
NVIDIA Riva with automatic speech recognition enabled
Why it's wrong here
Riva is a speech AI SDK for transcription and translation, not a text-based guardrail system. While it can process audio, it does not analyze LLM text outputs for PII or enforce content policies. Enabling ASR would be irrelevant to detecting patient names or medical record numbers in generated clinical documentation.
- ✗
NVIDIA TensorRT-LLM with quantization-aware training
Why it's wrong here
TensorRT-LLM is an inference optimization library that accelerates LLM execution through techniques like quantization and kernel fusion. It does not provide runtime content filtering or PII detection. Applying quantization-aware training might reduce model size but would not prevent the model from generating sensitive information in its outputs.
- ✗
NVIDIA Triton Inference Server with dynamic batching enabled
Why it's wrong here
Triton Inference Server is a model-serving platform that optimizes throughput and latency through features like dynamic batching, but it does not inspect or filter the content of LLM outputs. It has no built-in PII detection or guardrail capability, so it cannot block sensitive data leakage in generated text.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.