A team trains a transformer language model on a large corpus but the model achieves very low training loss while performing poorly on held-out text. Which action most directly addresses this outcome?
A large gap between near-zero training loss and weak held-out performance is the classic signature of overfitting. Dropout randomly deactivates units during training, preventing co-adaptation, while weight decay constrains parameter magnitude. Together they reduce memorization of training text and improve generalization to unseen sequences, directly targeting the observed symptom.
Why this answer
Low training loss combined with weak held-out performance indicates overfitting, meaning the model memorizes training text instead of learning generalizable patterns. Regularization techniques such as dropout and weight decay constrain model capacity and penalize reliance on specific training examples, which directly reduces the train-validation gap and improves performance on unseen text.
Exam trap
The trap here is reacting to poor validation results by training longer or tuning the learning rate, when the low training loss already proves the model is memorizing rather than underfitting.