Courseiva

CCNA Ncp Prompt Engineering Questions

36 questions · Ncp Prompt Engineering topic · All types, answers revealed

1
MCQmedium

An engineer is deploying an NVIDIA NeMo Guardrails system to moderate a chatbot's responses. The chatbot must refuse to answer questions about politics but should answer questions about weather. Which prompt engineering strategy in NeMo Guardrails is most appropriate to enforce this behavior?

A.Add a Python function that checks the user input for political keywords and returns a refusal, bypassing the model entirely.
B.Define a dialogue flow with a canonical form for political questions that triggers a refusal response, and a separate flow for weather questions that allows the answer.
C.Use a system prompt that instructs the model to refuse political questions and answer weather questions.
D.Fine-tune the underlying LLM to refuse political questions and answer weather questions.
AnswerB

NeMo Guardrails uses dialogue flows defined in Colang to specify how the bot should respond to different user intents. By creating a flow that detects political questions and triggers a refusal, and another flow for weather that permits answering, you enforce the desired behavior. This is the intended use of the guardrails framework.

Why this answer

NeMo Guardrails uses Colang dialogue flows to define bot behavior based on user intent. Creating separate flows for political and weather questions ensures deterministic refusal or answering. This is the core prompt engineering approach within the guardrails framework and provides reliable, auditable control.

Exam trap

The trap here is confusing runtime guardrail flows with model-level instructions or fine-tuning, which do not provide the same deterministic enforcement.

2
Multi-Selectmedium

A developer is creating prompts for an NVIDIA NIM-hosted LLM to summarize financial reports. The reports are lengthy and contain many tables and figures. The developer wants to ensure the summaries are accurate and include key numerical data. Which TWO prompt engineering techniques should be applied? (Choose two.)

Select 2 answers
A.Use a chain-of-thought prompt that asks the model to first list all tables and figures, then write the summary.
B.Set the temperature to a high value like 1.0 to encourage the model to creatively interpret the financial data and provide insightful analysis.
C.Instruct the model to extract and include all monetary values and percentages exactly as they appear in the report, and to avoid rounding or paraphrasing numbers.
D.Instruct the model to ignore any tables and figures and focus only on the narrative text to avoid confusion.
E.Provide a few-shot example of a summary that correctly includes key numbers from a sample report, demonstrating the desired format and level of detail.
AnswersC, E

This instruction directly addresses the need for accurate numerical data. By explicitly telling the model to include all monetary values and percentages exactly as they appear, and to avoid rounding or paraphrasing, you reduce the risk of the model altering numbers. It sets a clear constraint that the model must adhere to, which is crucial for financial summaries where precision is paramount.

Why this answer

To ensure accurate financial summaries with key numerical data, the developer should explicitly instruct the model to include all monetary values and percentages exactly as they appear, and provide a few-shot example demonstrating correct inclusion of numbers. These two techniques directly guide the model to preserve numerical accuracy and follow the desired format. High temperature, chain-of-thought, or ignoring tables would not achieve the goal.

Exam trap

The trap here is thinking that chain-of-thought or high temperature will improve numerical accuracy, when the real need is explicit instructions and examples that enforce exact inclusion of figures.

3
MCQeasy

A data scientist is using an NVIDIA NeMo LLM to generate Python code from natural language descriptions. The model often produces code that works but does not follow the team's style guide, such as using single quotes instead of double quotes and missing type hints. Which prompt engineering technique should the data scientist use to improve adherence to the style guide?

A.Chain-of-thought prompting to encourage the model to reason about the style guide before writing code.
B.Few-shot prompting with examples that demonstrate the desired coding style.
C.Zero-shot prompting with a detailed instruction describing the style guide rules.
D.Increasing the temperature to allow more creative code generation.
AnswerB

Few-shot prompting provides the model with concrete examples of the desired output style, such as double quotes and type hints. By including a few input-output pairs in the prompt, the model can infer the pattern and apply it to new queries. This is a standard prompt engineering technique to guide formatting and style without retraining.

Why this answer

Few-shot prompting with examples that embody the desired style guide is the most effective way to teach the model the specific formatting rules. By showing the model correct examples, it can mimic the style in new generations. This technique is particularly useful for coding tasks where precise syntax and style matter.

Exam trap

The trap here is assuming that a detailed zero-shot instruction is sufficient to enforce style, when in fact models often require concrete examples to reliably follow formatting rules.

4
MCQmedium

A developer is using an NVIDIA NIM for a customer support chatbot. The chatbot must handle multi-turn conversations and maintain context about the user's issue. The developer notices that after several turns, the bot starts giving generic responses and forgets earlier details. Which prompt engineering approach is most effective to maintain context?

A.Use a summarization prompt at each turn to condense the conversation so far, and include only the summary in the next prompt.
B.Add a system prompt that instructs the model to always ask the user to repeat their issue if it is unsure.
C.Use a higher temperature setting to make the model more creative in interpreting the user's issue.
D.Include the entire conversation history in the prompt for each turn, ensuring the model has access to all previous messages.
AnswerD

Including the full conversation history in the prompt allows the model to attend to all previous turns, maintaining context. This is a standard approach for multi-turn dialogues. However, it is limited by the context window size; if the conversation exceeds the window, older messages may be truncated. Despite this, for many support conversations, it is the most straightforward and effective method to preserve context without additional infrastructure.

Why this answer

Including the full conversation history in the prompt is the most direct way to maintain context across turns. The model can attend to all previous messages, ensuring it remembers details from earlier in the conversation. While summarization can reduce token usage, it risks losing critical information.

Temperature and system prompts do not address context retention.

Exam trap

The trap here is assuming that summarizing the conversation will preserve all necessary details, when in fact summarization can omit specifics that are crucial for support interactions.

5
Multi-Selecthard

A team is deploying an NVIDIA NIM for a Llama 3 model as a retrieval-augmented generation (RAG) assistant over internal documentation. Users report that the assistant sometimes answers from its pretrained knowledge instead of the retrieved passages, and occasionally cites a passage that does not support its claim. Which TWO prompt engineering changes best reduce these behaviors? (Choose two.)

Select 2 answers
A.Add a chain-of-thought instruction asking the model to reason about why the user's question is interesting before answering.
B.Raise the temperature to 0.8 so the model produces more varied answers and avoids memorized responses.
C.Instruct the model to answer only from the provided context and to respond with a fixed phrase when the context is insufficient.
D.Require the model to quote the exact sentence from the context that supports each claim before stating the answer.
E.Remove the retrieved passages from the prompt and rely on the model's pretrained knowledge to keep the prompt short.
AnswersC, D

An explicit grounding instruction tells the model that the retrieved passages are the sole source of truth. A fallback phrase for insufficient context prevents the model from filling gaps with pretrained knowledge. This directly addresses both symptoms: unsupported answers and reliance on internal memory, by defining what the model may use and what it must do when the context is inadequate.

Why this answer

Grounding the model with an explicit instruction to use only the provided context, plus a fallback for insufficient information, prevents reliance on pretrained knowledge. Requiring verbatim supporting quotes makes each claim auditable and discourages fabricated citations. Together these changes enforce source adherence and citation accuracy, while sampling changes or removing context would undermine the RAG design.

Exam trap

The trap here is assuming that increasing temperature or adding reasoning steps will improve faithfulness, when the real fix is constraining the model to the retrieved context and requiring verifiable quotes.

6
MCQhard

An engineer is using an NVIDIA NIM for a code generation model to produce Python functions from natural language descriptions. The model frequently generates code that uses deprecated libraries or incorrect function signatures. The engineer wants to improve the accuracy of the generated code by providing examples. Which prompting strategy is most appropriate?

A.Zero-shot prompting with a detailed description of the desired function and its parameters.
B.Few-shot prompting with examples that demonstrate the correct use of the desired libraries and function signatures.
C.Self-consistency prompting where the model generates multiple code solutions and the most common one is selected.
D.Chain-of-thought prompting that asks the model to first explain its reasoning about which libraries to use, then write the code.
AnswerB

Few-shot prompting provides the model with concrete examples of the desired output, including correct library usage and function signatures. By showing several input-output pairs where the output uses the correct libraries and signatures, the model can learn the pattern and apply it to new inputs. This is especially effective for code generation where precise syntax and API usage matter, and it directly addresses the deprecated library and signature issues.

Why this answer

Few-shot prompting with examples that demonstrate correct library usage and function signatures is the most direct way to guide the model to produce accurate code. By showing the desired pattern, the model can imitate it, reducing the likelihood of deprecated libraries or incorrect signatures. Other strategies like zero-shot, chain-of-thought, or self-consistency do not provide the concrete examples needed to correct systematic API errors.

Exam trap

The trap here is assuming that chain-of-thought or self-consistency will fix incorrect API usage, when the model needs concrete examples of correct code to learn the pattern.

7
Multi-Selecthard

A team is designing prompts for an NVIDIA NIM-hosted LLM that must produce concise, citation-backed answers from retrieved documents. They want to improve factual grounding and reduce unsupported claims. Which two prompt engineering practices best support this goal? (Choose two.)

Select 2 answers
A.Remove document identifiers from the context to simplify the prompt and reduce tokens.
B.Instruct the model to cite the specific document ID or snippet for each claim it makes.
C.Increase temperature to encourage the model to synthesize multiple documents creatively.
D.Instruct the model to say 'Insufficient evidence' when the retrieved documents do not contain the answer.
E.Allow the model to answer from general knowledge when retrieved documents are incomplete.
AnswersB, D

Requiring citations forces the model to tie statements to retrieved sources, making unsupported claims easier to detect and reducing free-form invention. It also gives reviewers a way to verify answers. This practice directly supports factual grounding because the model must reference evidence rather than rely on parametric memory.

Why this answer

Factual grounding in retrieval-augmented generation improves when the model must cite sources and when it is allowed to refuse when evidence is missing. Citations create traceability, and a refusal fallback prevents guessing. Allowing general knowledge, raising temperature, or stripping identifiers weakens provenance and increases unsupported claims.

Exam trap

The trap here is treating creativity or general knowledge as helpful, when citation-backed grounding requires restricting answers to retrieved evidence and permitting refusal.

8
MCQhard

A team fine-tunes an NVIDIA NeMo model to classify support tickets into five categories. In production, the model sometimes outputs free-form explanations instead of a single category label, breaking the downstream parser. Which prompt engineering change MOST reliably constrains the output format?

A.Add a polite request asking the model to 'try to keep answers short and to the point' while leaving the output format unspecified.
B.Increase max_tokens so the model has more room to explain its reasoning before stating the final category label.
C.Lower the temperature to zero and trust that deterministic sampling will force the model to emit only a category label.
D.Add three few-shot examples that each end with the exact category label, and instruct the model to output only the label with no additional text.
AnswerD

Few-shot examples demonstrate the precise output pattern, and the explicit instruction forbids extra text. Because the model conditions on the demonstrated format, it strongly biases toward emitting only a label. This combination is the most reliable prompt-level method to enforce a strict output contract without changing decoding or adding post-processing.

Why this answer

Strict output contracts are best enforced by showing the exact desired format through few-shot examples and explicitly prohibiting any additional text. Demonstrations act as in-context conditioning that shapes the model's continuation pattern, while the instruction closes the loophole of adding commentary. Sampling parameters control randomness, not structure.

Exam trap

The trap here is believing that temperature zero guarantees a clean label-only output, when format compliance is determined by prompt instructions and examples rather than by decoding settings.

9
MCQhard

A team uses an NVIDIA NIM-hosted model to draft release notes from a changelog. Reviewers report the drafts omit minor fixes and overstate the significance of small changes. Which prompt engineering adjustment BEST addresses both issues?

A.Instruct the model to include every changelog entry, to describe each change at its actual scope without exaggeration, and to organize the output by category such as features, fixes, and known issues.
B.Instruct the model to summarize only the three most impactful changes so the release notes stay concise and readable.
C.Add a single example of a polished release note and ask the model to 'write something similar' for the new changelog.
D.Increase the model's max_tokens and lower top_p so the output can be longer and more focused on the most probable phrasing.
AnswerA

Explicitly requiring complete coverage of all entries addresses omissions, while the accuracy instruction discourages inflating small changes. Categorizing output gives the model a structural checklist that makes it easier to verify that no entry was dropped and keeps minor fixes visible rather than buried.

Why this answer

The reported defects are content-level problems: missing entries and distorted emphasis. Only explicit instructions can fix them, by requiring complete coverage, accurate scoping of each change, and a categorized structure that makes omissions easy to spot. Sampling parameters and single stylistic examples influence form, not editorial completeness or proportionality.

Exam trap

The trap here is reaching for decoding parameters or a style example to fix what are actually content-coverage and accuracy problems that only explicit instructions can resolve.

10
Multi-Selecthard

When building an NVIDIA NeMo LLM application for automated document review, which THREE of the following prompt design choices are critical for ensuring high-quality output? (Select exactly THREE)

Select 3 answers
A.Use standardized delimiters to isolate user-provided documents.
B.Include instructions that prioritize speed over accuracy.
C.Require the output to be in a machine-readable format like JSON.
D.Instruct the model to perform a chain-of-thought validation of its findings.
E.Use a high temperature setting to ensure diverse review perspectives.
AnswersA, C, D

Properly isolating document segments prevents the model from conflating the input data with its own internal knowledge or instructions. This clarity is essential for document review applications where accuracy is paramount, as it ensures the model is specifically analyzing the provided text rather than hallucinating based on external training.

Why this answer

Effective document review requires accuracy, traceability, and consistency. Using clear delimitation prevents data corruption, requiring structured output (like JSON) allows for downstream programmatic integration, and chain-of-thought prompts ensure the model validates its conclusions against the text. These choices together create a robust, production-ready pipeline that minimizes errors and provides the necessary structure for automated workflows in enterprise environments.

Exam trap

Candidates often neglect the importance of machine-readable output formats, focusing only on the content of the response while ignoring the necessity of programmatic integration for automated document review workflows.

11
MCQmedium

Refer to the exhibit. Given this NeMo configuration, which prompt modification would best improve the reliability of technical support queries?

A.Append 'Be as creative as possible' to the system prompt.
B.Change the stop sequences to be empty.
C.Require the model to state 'I cannot answer this' if the manual is silent.
D.Increase the temperature to 0.9 to ensure varied answers.
AnswerC

This specific instruction provides a clear 'exit path' for the model when the provided context is inadequate. By explicitly defining the behavior for unsupported queries, you prevent the model from guessing or fabricating answers, which is crucial for maintaining the trust and reliability of your technical documentation bot.

Why this answer

The exhibit shows a relatively low temperature, which is good for consistency, but the system prompt lacks specific constraints on how to handle missing data. By modifying the prompt to strictly enforce grounding in the manuals and providing a fallback mechanism, you significantly improve the model's reliability in technical support scenarios. This is a critical step for maintaining quality in production environments utilizing NVIDIA NeMo technology.

Exam trap

Test-takers frequently choose to adjust model temperature or top-p sampling parameters instead of modifying the system prompt to explicitly handle missing data scenarios.

12
MCQmedium

A financial analyst is using an NVIDIA NIM-hosted Llama 3.1 70B model to extract key financial metrics from quarterly earnings call transcripts. The model inconsistently returns a prose summary instead of the required structured JSON. The analyst needs the output to be reliably parseable by a downstream script that expects a fixed schema with fields "revenue", "eps", and "guidance". Which prompt engineering technique is most appropriate to enforce this output format?

A.Increase the temperature parameter to 0.9 so the model explores more diverse output formats and may eventually produce JSON.
B.Use a chain-of-thought prompt asking the model to 'think step by step' before answering, which will naturally lead to JSON output.
C.Provide a JSON schema and a completed example in the prompt, and instruct the model to return only JSON matching that schema.
D.Add a system prompt that says 'You are a helpful assistant' and rely on the model's instruction-following ability to infer the JSON requirement.
AnswerC

This is correct because providing an explicit schema plus a worked example gives the model a concrete template to imitate, which strongly biases generation toward valid JSON. The instruction to return only JSON reduces the chance of prose leakage. This combination of structured format specification and demonstration is the most reliable way to enforce a fixed output schema without relying on post-processing.

Why this answer

Enforcing a structured output like JSON requires explicit specification of the schema and often a concrete example. Providing a schema and a completed example leverages the model's in-context learning to mimic the exact format, while the instruction to return only JSON reduces extraneous text. Other techniques like temperature adjustment or chain-of-thought do not directly control output structure and may even worsen format compliance.

Exam trap

The trap here is assuming that a generic instruction like 'return JSON' is sufficient without providing a schema or example, when models often need concrete demonstrations to reliably adhere to complex structured formats.

13
MCQeasy

A technical support team is building a chatbot using an NVIDIA NIM microservice. The chatbot must answer questions about a specific product's warranty policy. The team wants to ensure the model's responses are grounded in the official warranty document, which is 50 pages long, and avoid inventing policy details. Which prompt engineering approach is most effective for this scenario?

A.Ask the model to 'think step by step' and reason about the warranty policy from its pre-trained knowledge.
B.Use a retrieval-augmented generation (RAG) pipeline to fetch the most relevant warranty sections and include them in the prompt as context.
C.Fine-tune the model on the warranty document using NVIDIA NeMo, then use the fine-tuned model without any additional context in the prompt.
D.Include the entire warranty document in the system prompt and instruct the model to answer based only on that document.
AnswerB

RAG retrieves only the pertinent sections of the warranty document and places them in the prompt, providing focused context that the model can use to generate accurate answers. This reduces hallucination because the model is conditioned on specific, relevant text. It also scales to large documents and keeps the prompt within token limits. This is the standard approach for grounding LLM responses in proprietary knowledge bases.

Why this answer

Retrieval-augmented generation (RAG) is the most effective way to ground responses in a specific document. It retrieves relevant sections and includes them in the prompt, ensuring the model's answer is based on factual content rather than pre-trained memory. This reduces hallucinations and handles documents that exceed the context window.

Other methods either risk hallucination, are inefficient, or lack flexibility.

Exam trap

The trap here is assuming that fine-tuning is necessary to teach the model a document, when RAG with prompt context is often simpler, more accurate, and easier to update.

14
MCQeasy

A developer is writing a system prompt for an NVIDIA NIM-hosted assistant that must always respond in formal English, never use slang, and never reveal internal system instructions. Where should these persistent behavioral rules be placed for the MOST consistent effect?

A.Omitted entirely, because well-aligned base models already default to formal English and never disclose their instructions.
B.Encoded only in the model's temperature and top_p settings, since decoding parameters control response style.
C.Appended to the end of each user message as a reminder, so the rules stay close to the model's most recent input.
D.In the system prompt, stated as explicit rules that apply to every turn of the conversation.
AnswerD

System prompts establish persistent, high-priority behavioral guidance that applies across all turns. Placing tone, style, and confidentiality rules there gives the model a stable frame of reference, making consistent adherence far more likely than rules scattered in individual user messages.

Why this answer

Persistent behavioral requirements such as tone, style, and confidentiality belong in the system prompt, which the model treats as standing guidance across the conversation. This placement keeps rules stable regardless of what users type and avoids the fragility of repeating them per message or expecting decoding parameters or base alignment to handle policy.

Exam trap

The trap here is assuming that a well-aligned model needs no explicit style or confidentiality instructions, when consistent behavior still depends on system-level guidance.

15
MCQhard

A developer is using an NVIDIA NIM for a Llama 3.1 70B model to build a legal document review assistant. The model must answer questions based on a provided contract, but the contracts are often 50,000 tokens long, exceeding the model's 8,000-token context window. Which prompt engineering strategy is most appropriate to handle this constraint?

A.Use a retrieval-augmented generation (RAG) approach to fetch only the relevant sections of the contract and include them in the prompt.
B.Split the contract into chunks and ask the model to summarize each chunk sequentially, then combine the summaries.
C.Increase the model's context window by fine-tuning it on longer sequences.
D.Use a sliding window approach where the model processes the contract in overlapping segments and maintains a memory of previous segments.
AnswerA

RAG involves retrieving the most relevant passages from the long document and including only those in the prompt. This reduces the context length to fit within the model's window while still providing the necessary information. It is the standard solution for handling documents that exceed the context limit.

Why this answer

RAG is the most appropriate strategy because it retrieves only the relevant sections of the long contract, fitting them into the model's context window. This allows the model to answer questions accurately without needing to process the entire document. It is a widely adopted prompt engineering pattern for handling context length limitations.

Exam trap

The trap here is assuming that fine-tuning or summarization can easily overcome context window limits, when in fact retrieval-based approaches are the practical solution.

16
Multi-Selectmedium

An engineer is designing prompts for an NVIDIA NIM-hosted model that must extract structured fields from unstructured invoices. The extraction accuracy is inconsistent across vendors with different layouts. Which TWO prompt engineering techniques would MOST improve reliability? (Choose two.)

Select 2 answers
A.Instruct the model to output fields in a fixed JSON schema and to use null for any field not present in the document.
B.Ask the model to summarize the invoice in prose first and then extract fields from its own summary in a second pass.
C.Raise the temperature to 0.9 so the model explores multiple interpretations of ambiguous invoice text before committing to values.
D.Provide few-shot examples that cover several distinct invoice layouts, each showing the exact input-to-output field mapping.
E.Remove all formatting and concatenate the entire invoice text into a single lowercase string before sending it to the model.
AnswersA, D

A fixed schema removes ambiguity about output shape and key names, and the null convention gives the model a defined behavior for missing data instead of inventing values. Together these constraints make outputs parseable and reduce hallucinated field values, which is essential for downstream automation.

Why this answer

Reliable structured extraction combines format constraints with demonstrated patterns. A fixed JSON schema with a null convention makes outputs machine-parseable and defines behavior for absent fields, while few-shot examples across varied layouts teach the model to handle layout diversity. Together they reduce both structural errors and hallucinated values, which prose summarization and high-temperature sampling would worsen.

Exam trap

The trap here is treating higher temperature as a way to reason through ambiguity, when in extraction tasks it mainly adds run-to-run inconsistency and invented values.

17
MCQmedium

A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3.1 8B Instruct model. The assistant must answer questions about an order solely based on a JSON payload containing order details, and it must not use any outside knowledge. Which prompt engineering approach best ensures the model adheres to this constraint?

A.Prefix the prompt with a system message that says 'You are a helpful assistant' and then ask the question without including the JSON.
B.Include the JSON payload in the prompt and instruct the model to answer using only the provided data, adding a fallback phrase for missing information.
C.Fine-tune the model on a dataset of order-related questions and answers before deploying it as a NIM microservice.
D.Use a low temperature setting (e.g., 0.1) to make the model more deterministic and less likely to hallucinate.
AnswerB

This approach explicitly grounds the model in the provided JSON and sets a clear boundary: answer only from the data. Adding a fallback phrase for missing information prevents the model from inventing details, which is critical for factual accuracy in customer support. It leverages the instruction-following capability of Llama 3.1 without requiring additional guardrail systems.

Why this answer

Grounding the model with the exact JSON payload and instructing it to answer only from that data, with a fallback for missing information, ensures responses are based solely on the provided order details. This leverages the instruction-following ability of Llama 3.1 and avoids reliance on external knowledge. Other methods like fine-tuning or temperature adjustment do not enforce strict data adherence at inference time.

Exam trap

The trap here is assuming that a low temperature setting alone can prevent hallucination and enforce data grounding, when in fact explicit instruction and context inclusion are required.

18
MCQeasy

A developer is building an interactive assistant using NVIDIA NIM microservices. The assistant must answer questions about a specific set of internal policies. The developer wants to ensure the model's responses are grounded in those policies and not in its general pre-training knowledge. Which prompt engineering technique should be applied?

A.Fine-tune the model on the internal policy documents, then use a simple zero-shot prompt to ask questions.
B.Use a chain-of-thought prompt that asks the model to reason step by step about the policy before answering.
C.Increase the temperature setting to 0.9 to encourage the model to explore a wider range of policy interpretations.
D.Provide the relevant policy excerpts directly within the prompt, instruct the model to answer only from that context, and to state when the answer is not present.
AnswerD

This is retrieval-augmented generation (RAG) at the prompt level. By injecting the policy excerpts into the context and explicitly constraining the model to use only that information, you prevent it from relying on its pre-trained knowledge. Instructing it to say when the answer is missing further reduces hallucination and keeps the response grounded in the provided internal policies.

Why this answer

Grounding the model in a specific set of documents requires placing the relevant excerpts in the prompt and explicitly instructing the model to answer only from that context. This constrains the model to the provided facts and reduces hallucination. Other techniques like increasing temperature, chain-of-thought, or fine-tuning do not achieve this grounding for a dynamic, interactive policy assistant.

Exam trap

The trap here is assuming that fine-tuning or chain-of-thought alone will make the model answer from a specific set of documents, when in fact the documents must be supplied in the prompt context.

19
MCQmedium

When implementing Chain-of-Thought (CoT) prompting for a complex NVIDIA NeMo-based reasoning task, what is the primary benefit of encouraging the model to generate intermediate steps?

A.It forces the model to use more GPU memory per token.
B.It increases the likelihood of the model selecting a random seed.
C.It decomposes complex problems into verifiable logical segments.
D.It eliminates the need for system-level instructions entirely.
AnswerC

Breaking down multi-step problems into smaller, sequential steps allows the model to maintain context and reduces cumulative error rates. By generating intermediate proofs or calculations, the model provides a trace that developers can analyze to identify where the reasoning failed, which is vital for robust application development.

Why this answer

Chain-of-Thought prompting decomposes complex problems into sequential logical steps, which is critical when using LLMs for technical reasoning tasks. By forcing the model to articulate its internal logic, the likelihood of hallucination decreases significantly. This practice is essential for NVIDIA engineers deploying agents that require high-precision output, as it creates an audit trail for the model's reasoning process and allows for better debugging of multi-turn interactions.

Exam trap

Candidates often confuse CoT with simple 'few-shot' prompting, assuming the primary goal is just to provide examples rather than forcing the model to articulate the logical steps required for complex verification.

20
MCQmedium

A team is using an NVIDIA NIM for a Mistral model to classify support tickets into one of five fixed categories. Accuracy is inconsistent, and the model sometimes invents new categories. Which prompt engineering change is most likely to improve reliability without retraining the model?

A.Add a chain-of-thought instruction that asks the model to reason step by step before naming a category.
B.Increase the context window by concatenating the entire ticket history for every request.
C.Provide three labeled examples per category in the prompt and instruct the model to output only the category name.
D.Lower the temperature to 0.0 and remove all instructions so the model relies on its pretrained knowledge.
AnswerC

Few-shot examples define the exact label set and demonstrate the mapping from ticket text to category. Instructing the model to output only the category name removes room for invented labels. This directly addresses the inconsistency and the hallucinated categories without any fine-tuning, making it the most effective change for a fixed-label classification task.

Why this answer

Few-shot examples with explicit labels define the allowed output space, and the instruction to emit only the category name prevents invented labels. This combination directly targets both symptoms, inconsistency and hallucinated categories, and requires no model retraining, unlike sampling or reasoning changes that leave the label set undefined.

Exam trap

The trap here is believing that lowering temperature guarantees correct classification, when the real issue is that the allowed labels were never defined in the prompt.

21
MCQeasy

A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3 model. The assistant must always respond in valid JSON with keys 'category' and 'urgency'. The model often returns conversational text instead. Which prompt engineering change most directly enforces the required output format?

A.Append 'Return only JSON with keys category and urgency' and set the NIM request parameter 'guided_json' to the target schema.
B.Add 'Do not hallucinate' to the prompt and lower the max_tokens parameter to 50.
C.Add a system prompt that says 'You are a helpful assistant' and increase the temperature to 0.9.
D.Use few-shot examples of JSON outputs and set top_p to 0.1 without any schema constraint.
AnswerA

Combining an explicit instruction with NVIDIA NIM's guided_json parameter constrains decoding to the provided JSON schema, so the model cannot emit conversational text. The instruction aligns the model's intent while guided_json enforces structural validity at generation time. This is the most direct way to guarantee the required keys and format in the response.

Why this answer

Structured output requires both a clear instruction and a decoding-time constraint. NVIDIA NIM supports guided_json, which restricts token generation to a supplied JSON schema, guaranteeing valid keys and syntax. Pairing that with an explicit prompt instruction aligns the model's behavior with the schema.

Few-shot examples or temperature changes alone cannot guarantee strict JSON compliance.

Exam trap

The trap here is assuming that simply asking the model for JSON or providing examples is enough, when only schema-guided decoding like guided_json can guarantee valid structure.

22
MCQmedium

Which technique is most appropriate for a task requiring an LLM to generate code in a specific enterprise-internal syntax that is not well-represented in its public training data?

A.Zero-shot prompting with broad general coding instructions.
B.Few-shot prompting with multiple code examples.
C.Increasing the model temperature to encourage exploration.
D.Reducing the context window to force brevity.
AnswerB

By providing multiple examples of the target syntax, you enable the model to perform in-context learning of the specific patterns required. This pattern-matching approach allows the model to generalize the internal syntax correctly, providing accurate outputs that conform to enterprise standards without needing to perform full model retraining.

Why this answer

When dealing with proprietary or rare syntax, few-shot prompting with high-quality, representative examples is the most effective way to guide the model. By including these examples within the prompt, you provide the context the model lacks, significantly reducing syntax errors. This is a critical skill for NVIDIA developers building custom coding assistants for internal proprietary frameworks or legacy infrastructure support.

Exam trap

Candidates often mistakenly suggest fine-tuning as the first step, ignoring that few-shot prompting is a faster, more effective way to introduce specific, rare syntax without the overhead of retraining.

23
MCQmedium

An engineer is building a customer-facing FAQ bot using an NVIDIA NIM-hosted Llama 3.1 70B model. The bot must answer ONLY from a supplied product knowledge base and must respond with 'I don't have that information' when the answer is not present. Which prompt engineering approach BEST enforces this constraint?

A.Instruct the model to answer only from the provided knowledge base, explicitly state that it must reply with the refusal phrase when the answer is absent, and include the knowledge base inside clearly delimited sections.
B.Set the top_p value to 0.1 and rely on the model's pretrained knowledge of the product domain instead of supplying the knowledge base in the prompt.
C.Append the entire knowledge base to every user message without any instruction about scope or refusal behavior, letting the model infer the rules from context.
D.Increase the temperature to 1.0 and add 'Be creative and helpful' to the system prompt so the model can improvise when the knowledge base is incomplete.
AnswerA

Combining an explicit grounding instruction, a mandated refusal string, and clearly delimited context gives the model a precise behavioral contract. The delimiters separate trusted context from user input, while the refusal phrase defines the exact fallback behavior. This is the most reliable prompt-level method to constrain an NIM-hosted model to grounded answers without retraining.

Why this answer

Grounded FAQ bots need an explicit behavioral contract: what source to use, what to do when the source lacks the answer, and clear separation of context from user input. Delimiters prevent prompt injection from blending with instructions, and a fixed refusal phrase makes the fallback deterministic and testable. Sampling parameters alone cannot enforce scope constraints.

Exam trap

The trap here is assuming that lowering temperature or top_p will prevent hallucination, when grounding actually depends on explicit instructions and delimited context, not on sampling parameters.

24
MCQhard

An engineer is optimizing prompts for an NVIDIA NIM-hosted model used in a multi-turn technical troubleshooting chat. The model forgets earlier constraints, such as the customer's environment and the product version, as the conversation grows. Which prompt engineering technique best preserves these constraints across turns?

A.Repeat the full conversation history verbatim in every turn and increase max_tokens.
B.Maintain a compact running summary of key constraints and inject it into the system prompt each turn.
C.Use a single-turn prompt for each message and rely on the model's pretrained knowledge of the product.
D.Lower the temperature to 0 and remove the system prompt to avoid conflicting instructions.
AnswerB

A running summary keeps essential facts such as environment and product version in a stable, high-priority position, so they survive across turns without unbounded token growth. Injecting it into the system prompt gives those constraints consistent influence over every response. This directly addresses forgetting while controlling context length in a multi-turn chat.

Why this answer

Multi-turn memory is best preserved by extracting and re-injecting critical constraints rather than replaying all dialogue. A compact running summary placed in the system prompt keeps key facts salient and stable while avoiding context bloat. Temperature and single-turn designs do not address memory, and verbatim history can dilute attention and waste tokens.

Exam trap

The trap here is thinking that more conversation history or deterministic sampling will fix forgetting, when the real fix is summarizing and re-prioritizing the constraints each turn.

25
MCQmedium

A developer is using an NVIDIA NIM-hosted model to classify support tickets into a fixed set of categories. The model occasionally invents new category names. The team wants to guarantee that only allowed categories are returned. Which approach is most appropriate?

A.Provide the full category list in the prompt and use the NIM guided_choice parameter with those categories.
B.Use few-shot examples of each category and set top_k to 100.
C.Ask the model to explain its reasoning before choosing a category.
D.Add 'Choose from the list' to the prompt and set temperature to 1.0.
AnswerA

Listing the categories gives the model the allowed set, and guided_choice constrains decoding to exactly one of those strings. This eliminates invented labels at the sampling level, guaranteeing compliance. It is the most reliable method when the output must be one of a fixed enumeration.

Why this answer

When output must be one of a fixed set, enumerating the categories in the prompt and applying a decoding constraint such as guided_choice ensures only valid labels are emitted. Few-shot examples and reasoning improve quality but do not enforce the boundary. High temperature and broad top_k work against the requirement by increasing variability.

Exam trap

The trap here is relying on prompt wording or examples to restrict categories, when only an explicit enumeration combined with constrained decoding guarantees the model cannot invent a label.

26
MCQhard

Which of the following describes the 'Chain-of-Verification' (CoVe) prompting technique?

A.A method to verify that the GPU is running at full capacity.
B.A process where the model checks its own claims for factual consistency.
C.A technique to optimize the model's weights during training.
D.A protocol for securing the prompt against injection attacks.
AnswerB

CoVe explicitly mandates that the model critiques its own output. By drafting verification questions and answering them, the model can identify and correct errors in its initial draft. This self-correction loop is a powerful tool for improving the truthfulness and reliability of complex LLM-generated reports in enterprise environments.

Why this answer

Chain-of-Verification is a sophisticated technique designed to reduce hallucinations. It prompts the model to generate a response, then draft questions to verify its own claims, answer those questions, and finally revise the original response based on the verification. This is highly effective for NVIDIA developers building high-stakes applications where factual accuracy is non-negotiable and automated auditing is required.

Exam trap

Candidates often conflate CoVe with standard CoT, failing to realize that CoVe is specifically a post-generation verification loop aimed at fact-checking, rather than just a step-by-step reasoning process.

27
MCQmedium

Which prompt engineering strategy helps the model maintain focus when processing an extremely long document within a single context window?

A.Always increase the model's temperature to 1.0.
B.Instruct the model to analyze the document in sections and summarize each.
C.Remove all system-level instructions to save tokens.
D.Limit the context window size to 1024 tokens.
AnswerB

Breaking down a large document into segments for analysis ensures that the model provides equal attention to all parts. By summarizing each section before synthesis, you force the model to retain the key information from throughout the text, effectively mitigating the risk of ignoring information in the middle.

Why this answer

Long context windows can lead to the 'Lost in the Middle' phenomenon, where the model performs better on the beginning and end of the document but ignores the middle. Using 'attention-focusing' instructions or prompting the model to summarize segments before synthesizing a final answer helps maintain performance across the entire document. This is critical for NVIDIA engineers working with large-scale technical whitepapers and documentation.

Exam trap

Many candidates suggest increasing the context window size further, overlooking the 'Lost in the Middle' phenomenon where long documents are ignored.

28
MCQmedium

Refer to the exhibit. What prompt engineering strategy ensures the model consistently maintains its persona and technical expertise throughout this multi-turn dialogue?

A.Append the persona instruction to every user turn.
B.Use a persistent system prompt that defines the persona and constraints.
C.Increase the temperature to 1.0 to keep the model 'active'.
D.Force the model to provide a summary of its own persona at the end of each turn.
AnswerB

A well-defined system prompt acts as the 'source of truth' that the model refers back to in every turn. By embedding the persona and constraints (like 'CUDA expert') in the system prompt, you ensure that the model stays within its defined boundaries, regardless of the complexity of the interaction.

Why this answer

In multi-turn conversations, the model can 'drift' away from its original instructions. Re-affirming constraints or using a 'System Prompt' that is explicitly included in the context of every turn is critical. For an NVIDIA code optimizer, the model must maintain its technical persona, specifically focusing on CUDA performance, regardless of how complex the dialogue becomes, ensuring the expertise level remains consistent throughout.

Exam trap

Candidates assume a single initial prompt persists automatically across long multi-turn conversations without requiring reinforcement in subsequent turns.

29
MCQmedium

A team is using an NVIDIA NeMo-based LLM to answer questions over a product manual. The model sometimes answers using general knowledge instead of the provided manual excerpts. They want to force the model to rely only on the supplied context. Which prompt engineering approach best addresses this?

A.Instruct the model to answer only from the provided context and to reply 'Not in the provided context' when the answer is absent.
B.Increase the model's temperature so it explores more diverse answers from its pretrained knowledge.
C.Add more few-shot examples of correct answers without changing the instruction about using the context.
D.Shorten the prompt by removing the manual excerpts and rely on the model's product knowledge.
AnswerA

An explicit grounding instruction with a fallback phrase constrains the model to the supplied excerpts and gives it a safe response when the context lacks the answer. This reduces reliance on pretrained knowledge and makes unsupported answers visible. It is the most direct prompt-level fix for context adherence in a retrieval-augmented setup.

Why this answer

Grounding a model in retrieved context requires an explicit instruction that restricts answers to that context and defines what to do when the answer is missing. A fallback phrase such as 'Not in the provided context' prevents the model from filling gaps with pretrained knowledge. Few-shot examples help style but do not enforce the boundary; temperature and context removal work against the goal.

Exam trap

The trap here is believing that adding examples or tweaking sampling parameters will stop the model from using pretrained knowledge, when only an explicit grounding instruction with a refusal fallback reliably restricts it.

30
MCQhard

An engineer is using an NVIDIA NIM for a Mixtral model to extract structured data from invoices. The model occasionally returns fields with the wrong data type, such as a numeric amount as a string. The team wants a prompt engineering fix that does not require changing the model or adding a separate parser. Which approach is most effective?

A.Instruct the model to output the data in a natural language paragraph and then extract the fields manually.
B.Ask the model to double-check its output for type correctness before returning it.
C.Lower the temperature to 0.0 and increase the repetition penalty to discourage type mistakes.
D.Describe the desired output schema in the system prompt, including field names and expected data types, and provide one fully formatted example.
AnswerD

Specifying the schema with explicit data types and showing a complete example gives the model a precise template to follow. The example demonstrates the exact format, including numeric values without quotes. This combination is the most effective prompt-level fix because it removes ambiguity about both field names and types, reducing type errors without external parsing.

Why this answer

An explicit schema with field names and data types, paired with a complete example, gives the model an unambiguous template. The example shows numeric values without quotes, so the model imitates the correct types. Self-checking, sampling parameters, and natural language output do not provide the concrete structural guidance needed to eliminate type errors at the prompt level.

Exam trap

The trap here is relying on self-checking or sampling parameters to fix data type errors, when the prompt never defined the expected types in the first place.

31
MCQeasy

What is the primary purpose of 'Few-Shot Prompting' in the context of LLM optimization?

A.To reduce the latency of the underlying GPU cluster.
B.To provide in-context learning examples to guide output.
C.To compress the model weights for deployment.
D.To permanently store data in the model's internal memory.
AnswerB

Providing examples allows the model to observe the desired pattern of input and output. This pattern-matching capability enables the model to perform new tasks accurately without formal retraining, making it an ideal strategy for quickly adapting pre-trained models to specific enterprise data formats and business logic requirements.

Why this answer

Few-shot prompting involves providing a few examples of input-output pairs within the prompt to guide the model's performance on a specific task. This approach helps the model learn the desired format, tone, and logic without the need for intensive fine-tuning. It is a highly efficient way to steer model behavior for specific enterprise use cases, ensuring consistent results across multiple interaction sessions.

Exam trap

Candidates often confuse few-shot prompting with fine-tuning, assuming the model's weights are updated during the few-shot process rather than just providing context within the prompt.

32
MCQmedium

A developer is prompting an NVIDIA NIM for a Code Llama model to generate a Python function. The model produces correct logic but frequently omits type hints and docstrings, which the team requires. Which prompting technique best addresses this specific gap?

A.Include a short example of the desired function signature, type hints, and docstring in the prompt, then ask for the new function in the same style.
B.Increase max_tokens so the model has more room to include type hints and docstrings.
C.Add the instruction "write clean, production-quality code" to the system message.
D.Ask the model to first explain its reasoning about the function, then output the code.
AnswerA

Providing a concrete example that exhibits type hints and a docstring demonstrates the exact format expected. The model imitates the pattern it sees, so this one-shot demonstration directly fills the missing elements. It is more reliable than vague style instructions because the required structure is shown rather than described, leaving little room for interpretation.

Why this answer

A concrete example showing the desired function signature, type hints, and docstring teaches the model the exact format by imitation. Vague quality instructions, larger token budgets, and reasoning steps do not specify these structural requirements, so they fail to close the gap. The example-based approach is the most direct and reliable fix.

Exam trap

The trap here is using subjective quality phrases like production-quality code instead of demonstrating the exact structural elements the model must include.

33
MCQhard

Refer to the exhibit. Which prompt engineering technique would best force the model to prioritize technical detail over marketing language?

A.Add a constraint: 'Exclude marketing language and include specific architectural metrics like TDP, memory bandwidth, and interconnect speeds.'
B.Ask the model to 'Write a shorter summary' in the user prompt.
C.Decrease the temperature to 0.0 to make the model more factual.
D.Use few-shot prompting with generic summaries.
AnswerA

This approach provides clear negative constraints (exclude marketing) and positive constraints (include specific metrics). By defining the required output format and content, the model is compelled to ignore its tendency to generate generic, flowery text and focus on the hard data points that define the technical architecture.

Why this answer

The model is failing to adhere to the implicit expectation of 'technical depth.' By explicitly defining a structural constraint—such as requiring specific architectural metrics or removing marketing terminology—the model is forced to prioritize the technical aspects requested. This type of constraint-driven prompting is essential for professional NVIDIA technical writers and engineers who need precise documentation summaries.

Exam trap

Candidates frequently choose vague instructions like 'be more technical' instead of providing specific, actionable constraints, failing to realize that LLMs require explicit structural boundaries to filter out marketing-heavy language.

34
MCQmedium

A team is using an NVIDIA NIM-hosted Llama model to generate product descriptions from a list of technical specifications. The descriptions sometimes omit key specifications or include invented features. The team wants to improve reliability without changing the model. Which prompt engineering change is most likely to reduce these errors?

A.Add a system prompt that instructs the model to act as a technical writer and to include every specification exactly as given, and provide a few-shot example of a correct description.
B.Increase the max_tokens parameter to allow the model to generate longer descriptions, ensuring all specifications are covered.
C.Use a chain-of-thought prompt that asks the model to list each specification and then write the description.
D.Set the temperature to 0.0 to make the model deterministic and reduce creativity.
AnswerA

A system prompt sets the model's role and constraints, and few-shot examples demonstrate the desired output format and level of detail. By showing a correct description that includes all specifications without additions, the model is more likely to follow that pattern. This directly addresses both omission and invention by providing clear instructions and a concrete example, without retraining the model.

Why this answer

To reduce omissions and inventions without changing the model, the most effective prompt engineering approach is to provide explicit instructions and few-shot examples. A system prompt that defines the role and constraints, combined with an example that demonstrates including all specifications and avoiding fabrication, guides the model to produce more reliable outputs. Other parameter changes or chain-of-thought do not directly address content fidelity.

Exam trap

The trap here is thinking that lowering temperature or increasing max_tokens will fix content errors, when the real solution is to explicitly instruct the model and show it an example of the desired output.

35
MCQeasy

An engineer is building a customer support assistant using an NVIDIA NIM for a Llama 3 70B model. The assistant must always respond in valid JSON containing exactly the keys "issue" and "urgency", and must never include any other text. Which prompt engineering approach most directly enforces this output contract?

A.Raise the temperature to 1.0 so the model explores more possible response formats and eventually produces valid JSON.
B.Append the phrase "Please be concise" to every user message and rely on the model to infer the JSON structure.
C.Set the top_p value to 0.1 and the max_tokens parameter to 50 to force the model into a JSON-only response mode.
D.Add a system message that specifies the JSON schema and instructs the model to output only JSON, then validate responses programmatically.
AnswerD

A system message defining the exact schema and the instruction to emit only JSON constrains the model's behavior at the highest-priority level of the prompt. Because the model sees this before the user turn, it shapes every completion. Programmatic validation then catches any residual drift, making this the most direct and reliable enforcement for a strict output contract.

Why this answer

The most direct way to enforce a strict JSON contract is to state the schema and the output-only requirement in the system message, which has the highest priority in the prompt hierarchy, and then validate the result programmatically. Sampling parameters and stylistic instructions do not constrain structure, so they cannot guarantee the two required keys or the absence of extra text.

Exam trap

The trap here is assuming that sampling parameters such as temperature or top_p can enforce an output format, when they only affect randomness and token selection.

36
MCQmedium

When evaluating an LLM's response to a complex prompt, what is the 'Persona Adoption' technique?

A.A method to verify if the model has memorized personal data.
B.A method to assign an expert identity to improve response quality.
C.A way to force the model to identify the user's persona.
D.A technique to reduce the model's context window usage.
AnswerB

By setting a persona, such as 'Senior NVIDIA GPU Architect,' the model is primed to utilize more relevant technical terminology and adopt a problem-solving approach consistent with that role. This significantly improves the quality and relevance of the response compared to a generic or default conversational persona.

Why this answer

Persona adoption involves assigning a specific role, expertise level, or professional identity to the LLM within the system prompt. This technique helps calibrate the tone, vocabulary, and depth of the response to match the user's expectations. For NVIDIA applications, this ensures that the model speaks with the authority and technical precision required for engineering and developer-facing communications.

Exam trap

Candidates confuse persona adoption with few-shot prompting or fine-tuning, thinking it requires training data rather than simple system prompt instructions.

Ready to test yourself?

Try a timed practice session using only Ncp Prompt Engineering questions.