NCP-GENL Prompt Engineering Practice Question
A team is deploying an NVIDIA NIM for a Llama 3 model as a retrieval-augmented generation (RAG) assistant over internal documentation. Users report that the assistant sometimes answers from its pretrained knowledge instead of the retrieved passages, and occasionally cites a passage that does not support its claim. Which TWO prompt engineering changes best reduce these behaviors? (Choose two.)
⚠ Common exam trap
The trap here is assuming that increasing temperature or adding reasoning steps will improve faithfulness, when the real fix is constraining the model to the retrieved context and requiring verifiable quotes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Instruct the model to answer only from the provided context and to respond with a fixed phrase when the context is insufficient.
Grounding the model with an explicit instruction to use only the provided context, plus a fallback for insufficient information, prevents reliance on pretrained knowledge. Requiring verbatim supporting quotes makes each claim auditable and discourages fabricated citations. Together these changes enforce source adherence and citation accuracy, while sampling changes or removing context would undermine the RAG design.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add a chain-of-thought instruction asking the model to reason about why the user's question is interesting before answering.
Why it's wrong here
Reasoning about the question's interest level does not tie the answer to the retrieved context. It adds tokens and may distract the model from the grounding task. Chain-of-thought can help with multi-step logic, but here the failure is source adherence, so this change does not reduce unsupported answers or incorrect citations.
- ✗
Raise the temperature to 0.8 so the model produces more varied answers and avoids memorized responses.
Why it's wrong here
Higher temperature increases randomness and creativity, which makes hallucination and unsupported claims more likely, not less. The problem is that the model is ignoring or misusing the provided context, and sampling diversity does not address grounding. In a RAG assistant, lower temperature is generally preferred to keep answers faithful to the retrieved passages.
- ✓
Instruct the model to answer only from the provided context and to respond with a fixed phrase when the context is insufficient.
Why this is correct
An explicit grounding instruction tells the model that the retrieved passages are the sole source of truth. A fallback phrase for insufficient context prevents the model from filling gaps with pretrained knowledge. This directly addresses both symptoms: unsupported answers and reliance on internal memory, by defining what the model may use and what it must do when the context is inadequate.
- ✓
Require the model to quote the exact sentence from the context that supports each claim before stating the answer.
Why this is correct
Forcing a verbatim quote creates a verifiable link between each claim and the retrieved text. If no supporting sentence exists, the model cannot fabricate a quote easily, which reduces unsupported citations. This technique also makes it easier for downstream reviewers or automated checks to confirm that the cited passage actually supports the answer.
- ✗
Remove the retrieved passages from the prompt and rely on the model's pretrained knowledge to keep the prompt short.
Why it's wrong here
Removing the retrieved passages eliminates the grounding source entirely, guaranteeing that the model answers from pretrained knowledge. This is the opposite of what a RAG assistant should do and would worsen both reported problems. The passages must remain in the prompt so the model has authoritative content to cite and summarize.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.