NCA-GENL Trustworthy AI Practice Question
An AI platform team is preparing an LLM for a public-facing legal information assistant. During evaluation, they observe that the model gives systematically different quality answers depending on the dialect used in the prompt. Which action most directly addresses this Trustworthy AI concern?
⚠ Common exam trap
The trap here is assuming a system prompt or lower temperature can equalize treatment, when dialect disparities are baked into training data and weights.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Curate and augment training and evaluation data to include balanced dialect representation, then re-measure quality per dialect.
Dialect-dependent quality differences are a fairness and bias issue rooted in data representation. The durable fix is to rebalance training and evaluation data across dialects and then verify with per-dialect metrics. Prompt instructions, decoding parameters, and output normalization do not alter the learned statistical disparities, so they cannot reliably eliminate the gap or demonstrate improvement to reviewers.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Curate and augment training and evaluation data to include balanced dialect representation, then re-measure quality per dialect.
Why this is correct
Quality disparities tied to dialect stem from imbalanced representation in training and evaluation data. By deliberately curating balanced dialect samples and measuring performance per dialect, the team can identify and reduce the gap at its source. Re-measurement closes the loop, ensuring the intervention is verified rather than assumed, which aligns with trustworthy AI evaluation practice.
- ✗
Restrict the assistant to answering only in a single standardized dialect.
Why it's wrong here
Forcing one dialect may appear to equalize treatment, but it penalizes users who write in other dialects and can degrade comprehension and accessibility. It also does not fix the model's internal disparities; it merely hides them behind a normalization layer. This approach conflicts with inclusive design principles and would likely raise new fairness concerns.
- ✗
Lower the model temperature to make responses more deterministic across all users.
Why it's wrong here
Temperature controls randomness in sampling, not the model's learned associations between dialect and answer quality. Deterministic decoding would simply produce the same biased outputs consistently. This option confuses output variability with fairness and does nothing to address the underlying representation imbalance causing the disparity.
- ✗
Add a system prompt instructing the model to treat all users equally.
Why it's wrong here
A system prompt can shape tone but does not change the underlying statistical disparities learned from training data. Dialect-correlated quality gaps originate in the data distribution and model weights, so an instruction alone will not reliably close them. This is a superficial mitigation that may mask the issue during spot checks while leaving the disparity intact for real users.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.