A data science team is preparing an instruction fine-tuning dataset in NVIDIA NeMo Framework. They notice that after training, the model performs well on the training instructions but poorly on paraphrased versions of the same instructions. They want to improve generalization without increasing dataset size. Which data preparation change is most appropriate?
Paraphrasing and template diversification expose the model to varied surface forms of the same intent, which teaches it to map meaning rather than exact wording. This directly improves generalization to unseen paraphrases without adding new intents, and it is a standard data augmentation technique for instruction tuning.
Why this answer
Generalization to paraphrased instructions depends on the model seeing the same intent expressed in varied surface forms. Paraphrasing and template diversification teach the model to respond to meaning rather than exact wording, directly improving performance on unseen paraphrases without enlarging the dataset or changing training hyperparameters.
Exam trap
The trap here is treating poor paraphrase generalization as an optimizer or epoch problem, when the real fix is increasing surface-form diversity in the instruction data itself.