NCP-GENL Prompt Engineering Practice Question
Exhibit
System: You are an NVIDIA GPU code optimizer. User: Optimize this kernel: [code snippet]. Assistant: [Generates optimized CUDA code] User: Explain why this is faster. Assistant: [Generates explanation] User: Can you try another approach using shared memory?
Refer to the exhibit. What prompt engineering strategy ensures the model consistently maintains its persona and technical expertise throughout this multi-turn dialogue?
⚠ Common exam trap
Candidates assume a single initial prompt persists automatically across long multi-turn conversations without requiring reinforcement in subsequent turns.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a persistent system prompt that defines the persona and constraints.
In multi-turn conversations, the model can 'drift' away from its original instructions. Re-affirming constraints or using a 'System Prompt' that is explicitly included in the context of every turn is critical. For an NVIDIA code optimizer, the model must maintain its technical persona, specifically focusing on CUDA performance, regardless of how complex the dialogue becomes, ensuring the expertise level remains consistent throughout.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Append the persona instruction to every user turn.
Why it's wrong here
Appending instructions to every user turn is redundant and wastes tokens. The system prompt should handle persona definition at the foundational level of the API. If the model is drifting, it indicates a need for a stronger, more persistent system prompt or better guardrail configuration, not user-level repetition.
- ✓
Use a persistent system prompt that defines the persona and constraints.
Why this is correct
A well-defined system prompt acts as the 'source of truth' that the model refers back to in every turn. By embedding the persona and constraints (like 'CUDA expert') in the system prompt, you ensure that the model stays within its defined boundaries, regardless of the complexity of the interaction.
- ✗
Increase the temperature to 1.0 to keep the model 'active'.
Why it's wrong here
Active engagement is not achieved through temperature. High temperature will cause the model to deviate from the expert persona, potentially generating non-technical or inaccurate code. To maintain a consistent expert persona, you need low-temperature, deterministic behavior that stays strictly within the technical boundaries you have established.
- ✗
Force the model to provide a summary of its own persona at the end of each turn.
Why it's wrong here
This is an inefficient use of tokens and creates a repetitive, annoying user experience. The model's persona should be implicit in its output quality, not explicitly stated like a mantra. If the model is not acting like an expert, the system prompt needs refinement, not a self-summary. It's a waste.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.