NCP-GENL Fine-Tuning Practice Question
A company wants to teach a pretrained LLM to follow a specific output format for customer support replies using supervised fine-tuning on NVIDIA GPUs. Which data preparation approach best matches supervised fine-tuning for instruction following?
⚠ Common exam trap
Many candidates confuse any training on domain text with instruction tuning, when only prompt-response pairs teach the model to follow a requested format.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Collect pairs of an instruction prompt and the desired response, then train the model to predict the response tokens
Instruction-following SFT learns from prompt and response pairs with the loss applied to response tokens, which directly teaches the desired output format. Continued pretraining on unlabeled text, RLHF with a reward model, and response-only training all lack the prompt-conditioned supervised signal that this task requires.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Score model outputs with a reward model and update the policy using a policy gradient objective
Why it's wrong here
Reward-model scoring with a policy gradient objective describes reinforcement learning from human feedback, not supervised fine-tuning. It requires a trained reward model and preference data, adding substantial complexity. The scenario asks for a direct supervised approach to teach a format, which is simpler and better matched to prompt-response pairs.
- ✗
Provide only the desired responses without prompts and train the model to reproduce them verbatim
Why it's wrong here
Training on responses without their prompts removes the conditioning signal the model needs to learn when a given format applies. The model would memorize response text rather than learn to follow instructions, and it would not generalize to new prompts. Prompt-response pairing is essential for instruction following.
- ✓
Collect pairs of an instruction prompt and the desired response, then train the model to predict the response tokens
Why this is correct
Supervised fine-tuning for instruction following uses prompt and response pairs where the loss is computed on the response tokens. This teaches the model the mapping from instruction to desired output format. NeMo's SFT data formats, such as the prompt-completion and chat schemas, are built around exactly this structure, making it the correct data preparation approach.
- ✗
Gather a large unlabeled corpus and train the model to predict the next token across all documents
Why it's wrong here
Predicting the next token on unlabeled text is continued pretraining, which adapts the model to a domain but does not teach a specific instruction-following output format. Without prompt and response structure, the model has no signal about which behavior is desired. This approach also requires far more data and compute than SFT.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.