A team is fine-tuning a large language model on custom data using Vertex AI. They find that the training loss decreases but validation loss increases. What is the best course of action?
Adding dropout regularisation directly counteracts the overfitting causing validation loss to diverge from training loss. Reducing model capacity likewise limits memorisation of the custom fine-tuning data. Both address the generalisation gap the stem describes, where training loss falls while validation loss rises, restoring alignment between the two curves.
Why this answer
The increasing validation loss while training loss decreases is a classic sign of overfitting, where the model memorizes the training data but fails to generalize. Reducing model size or adding dropout regularization directly combats overfitting by limiting the model's capacity or introducing noise during training, which forces the model to learn more robust features. This is the best course of action because it addresses the root cause without further exacerbating the problem.
Exam trap
Google Cloud often tests the distinction between underfitting and overfitting, and the trap here is that candidates may confuse increasing validation loss with underfitting and incorrectly choose to increase epochs or learning rate, rather than recognizing the hallmark divergence of overfitting.
How to eliminate wrong answers
Option A is wrong because increasing the number of training epochs would further overfit the model to the training data, worsening the validation loss. Option C is wrong because increasing the learning rate can cause the model to overshoot minima and destabilize training, potentially increasing both training and validation loss, and does not address overfitting. Option D is wrong because switching to a smaller batch size introduces more noise in gradient estimates, which can sometimes help generalization but is not a direct or reliable remedy for overfitting; it may also slow convergence and is not the primary solution for the described loss divergence.