Courseiva
Trustworthy AI →mediumMultiple Choice

NCA-GENL Trustworthy AI Practice Question

An AI team is deploying a Llama 3 70B model for internal knowledge retrieval. They want to ensure that the model's responses are grounded in the company's approved document corpus and that any attempt to elicit unapproved content is blocked. Which NVIDIA NeMo Guardrails component should they configure to define these behavioral constraints?

⚠ Common exam trap

Many exam-takers confuse retrieval augmentation or inference optimization with guardrail enforcement, assuming any component that touches the model can enforce content policies.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Colang flows that specify dialogue patterns and guardrail actions.

NeMo Guardrails uses Colang to define dialogue flows and guardrail actions that constrain LLM behavior. By writing Colang flows, the team can enforce grounding to approved documents and block attempts to elicit unapproved content. Other components like embedding models or inference servers do not provide this policy enforcement capability.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Colang flows that specify dialogue patterns and guardrail actions.

    Why this is correct

    Colang is the modeling language used to define conversational flows and guardrails in NeMo Guardrails. By writing Colang flows, the team can specify allowed and disallowed topics, enforce grounding to the approved corpus, and trigger actions like blocking or redirecting responses. This directly addresses the requirement to constrain behavior and prevent unapproved content.

  • ✗

    A custom embedding model fine-tuned on the document corpus.

    Why it's wrong here

    Fine-tuning an embedding model improves retrieval relevance but does not enforce behavioral constraints or block unapproved content. It cannot define dialogue patterns or guardrail actions. While it can help ground responses, it lacks the control logic to prevent jailbreaks or off-topic responses, which are central to the team's requirement.

  • ✗

    TensorRT-LLM optimization profiles for inference acceleration.

    Why it's wrong here

    TensorRT-LLM optimizes inference speed and throughput but has no role in defining or enforcing content policies. It does not provide mechanisms to ground responses or block disallowed content. Using it would improve performance, not trustworthiness constraints, so it fails to meet the stated objective.

  • ✗

    The NVIDIA Triton Inference Server model ensemble configuration.

    Why it's wrong here

    Triton ensembles orchestrate multiple models for inference but do not define conversational guardrails or content restrictions. They cannot specify allowed topics or block unapproved outputs. While useful for serving, they do not address the need for behavioral constraints and grounding to the approved corpus.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.