A data scientist is preparing a labeled dataset of 50,000 customer support tickets for supervised fine-tuning of an LLM. Each ticket must be assigned exactly one of eight department labels. Which loss function is most appropriate for training this classification head?
Categorical cross-entropy compares the predicted probability distribution over the eight mutually exclusive department labels with the one-hot true label, penalizing probability mass placed on incorrect classes. It is the standard objective for single-label multi-class classification and produces well-calibrated softmax outputs, making it the right choice when each ticket belongs to exactly one department.
Why this answer
Because every support ticket carries exactly one of eight mutually exclusive department labels, the task is single-label multi-class classification. Categorical cross-entropy, paired with a softmax output layer, directly maximizes the probability of the correct department while normalizing across all eight classes, giving the strongest and most stable training signal for this scenario.
Exam trap
The trap here is confusing single-label multi-class classification, which uses categorical cross-entropy, with multi-label tagging, which uses binary cross-entropy.