Courseiva
Fine-Tuning →easyMultiple Choice

NCP-GENL Fine-Tuning Practice Question

A company wants to teach a pretrained LLM to follow a specific output format for customer support replies using supervised fine-tuning on NVIDIA GPUs. Which data preparation approach best matches supervised fine-tuning for instruction following?

⚠ Common exam trap

Many candidates confuse any training on domain text with instruction tuning, when only prompt-response pairs teach the model to follow a requested format.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Collect pairs of an instruction prompt and the desired response, then train the model to predict the response tokens

Instruction-following SFT learns from prompt and response pairs with the loss applied to response tokens, which directly teaches the desired output format. Continued pretraining on unlabeled text, RLHF with a reward model, and response-only training all lack the prompt-conditioned supervised signal that this task requires.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Score model outputs with a reward model and update the policy using a policy gradient objective

    Why it's wrong here

    Reward-model scoring with a policy gradient objective describes reinforcement learning from human feedback, not supervised fine-tuning. It requires a trained reward model and preference data, adding substantial complexity. The scenario asks for a direct supervised approach to teach a format, which is simpler and better matched to prompt-response pairs.

  • ✗

    Provide only the desired responses without prompts and train the model to reproduce them verbatim

    Why it's wrong here

    Training on responses without their prompts removes the conditioning signal the model needs to learn when a given format applies. The model would memorize response text rather than learn to follow instructions, and it would not generalize to new prompts. Prompt-response pairing is essential for instruction following.

  • ✓

    Collect pairs of an instruction prompt and the desired response, then train the model to predict the response tokens

    Why this is correct

    Supervised fine-tuning for instruction following uses prompt and response pairs where the loss is computed on the response tokens. This teaches the model the mapping from instruction to desired output format. NeMo's SFT data formats, such as the prompt-completion and chat schemas, are built around exactly this structure, making it the correct data preparation approach.

  • ✗

    Gather a large unlabeled corpus and train the model to predict the next token across all documents

    Why it's wrong here

    Predicting the next token on unlabeled text is continued pretraining, which adapts the model to a domain but does not teach a specific instruction-following output format. Without prompt and response structure, the model has no signal about which behavior is desired. This approach also requires far more data and compute than SFT.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.