A team is using an NVIDIA NIM for a Mistral model to classify support tickets into one of five fixed categories. Accuracy is inconsistent, and the model sometimes invents new categories. Which prompt engineering change is most likely to improve reliability without retraining the model?
Few-shot examples define the exact label set and demonstrate the mapping from ticket text to category. Instructing the model to output only the category name removes room for invented labels. This directly addresses the inconsistency and the hallucinated categories without any fine-tuning, making it the most effective change for a fixed-label classification task.
Why this answer
Few-shot examples with explicit labels define the allowed output space, and the instruction to emit only the category name prevents invented labels. This combination directly targets both symptoms, inconsistency and hallucinated categories, and requires no model retraining, unlike sampling or reasoning changes that leave the label set undefined.
Exam trap
The trap here is believing that lowering temperature guarantees correct classification, when the real issue is that the allowed labels were never defined in the prompt.