NCA-GENL Software Development Practice Question
A developer is preparing a RAG service that calls an NVIDIA-hosted LLM endpoint and must reduce hallucinations for questions whose answers are absent from the retrieved context. Which two practices should be applied in the application layer? (Choose two.)
⚠ Common exam trap
The trap here is assuming that more retrieved context or higher sampling temperature improves answer quality, when both actually increase the chance of unsupported output.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Instruct the model in the system prompt to answer only from the supplied context and to state explicitly when the context is insufficient.
Hallucination in RAG is controlled by constraining generation and by refusing to generate when retrieval fails. A grounding system prompt gives the model permission to abstain and limits it to supplied evidence, while a relevance threshold stops irrelevant chunks from ever reaching the prompt. Together they address both the model's behavior and the quality of its input.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Instruct the model in the system prompt to answer only from the supplied context and to state explicitly when the context is insufficient.
Why this is correct
A grounding instruction constrains generation to the retrieved passages and gives the model a sanctioned way to decline, which is the cheapest and most direct hallucination control. It works with any hosted endpoint because it lives entirely in the request payload. Combined with a threshold on retrieval scores, it prevents the model from inventing an answer when nothing relevant was retrieved.
- ✗
Increase the number of retrieved chunks to the model's full context limit so the answer is guaranteed to be somewhere in the prompt.
Why it's wrong here
Filling the context window with loosely related chunks dilutes the relevant evidence and can push the model toward spurious connections, while also inflating prefill cost and latency. Recall is not the problem being solved; precision is. More context without a relevance filter raises, rather than lowers, the chance of an unsupported answer.
- ✗
Cache every generated response and replay it for semantically similar questions to keep the answers consistent over time.
Why it's wrong here
Response caching improves latency and cost, and consistency is not the same as correctness. Replaying a cached answer for a merely similar question can return a confident response that does not match the current retrieved evidence. Caching does nothing to prevent hallucination for questions the cache has not seen, which is the scenario under test.
- ✓
Apply a minimum relevance score to retrieved chunks and skip generation entirely when no chunk clears the threshold.
Why this is correct
A retrieval threshold turns an unanswerable query into an explicit abstention instead of a prompt stuffed with irrelevant text. Irrelevant context is a leading cause of fabricated answers because the model tries to reconcile passages that do not address the question. Gating generation on relevance is a deterministic guard that complements prompt-level grounding instructions.
- ✗
Set the sampling temperature to its maximum so the model explores a wider range of candidate answers and avoids repeating a single phrasing.
Why it's wrong here
High temperature increases randomness precisely where determinism matters, making the model more likely to produce fluent but unsupported statements. For grounded question answering, low temperature is preferred so the model adheres to the retrieved evidence. Raising it works against the goal of reducing hallucination.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.