Refer to the exhibit. Which adjustment should you make to the configuration if the model shows signs of overfitting on a small, domain-specific dataset?
Exhibit
Config: { "lora_r": 64, "lora_alpha": 16, "target_modules": ["q_proj", "v_proj"], "task_type": "CAUSAL_LM" }Trap 1: Increase lora_r to 128
Increasing the rank parameter expands the capacity of the adapter matrices, which likely exacerbates overfitting. When a model overfits, it has already captured too much noise from the training set; providing more trainable parameters allows the model to memorize further, worsening the generalization gap rather than solving the issue.
Trap 2: Change target_modules to include all linear layers
Expanding the target modules increases the total count of trainable parameters, which typically increases the likelihood of overfitting on small datasets. For high-quality fine-tuning on limited data, fewer targeted modules are usually preferred to maintain regularization and focus the training on the most impactful attention mechanisms within the model.
Trap 3: Increase lora_alpha to 64
Increasing the alpha parameter effectively increases the weight of the LoRA adapters relative to the pre-trained weights. This makes the model more sensitive to the fine-tuning data, which would likely worsen overfitting rather than improving generalization. Alpha should generally be reduced or kept proportional to rank for stability.
- A
Increase lora_r to 128
Why it fails: Increasing the rank parameter expands the capacity of the adapter matrices, which likely exacerbates overfitting. When a model overfits, it has already captured too much noise from the training set; providing more trainable parameters allows the model to memorize further, worsening the generalization gap rather than solving the issue.
- B
Decrease lora_r to 8
Reducing the rank parameter restricts the model's ability to store specific details from the training data, effectively functioning as a form of regularization. By using a lower rank, the model is forced to focus on the most salient features of the dataset, which helps mitigate overfitting on smaller, specialized training sets.
- C
Change target_modules to include all linear layers
Why it fails: Expanding the target modules increases the total count of trainable parameters, which typically increases the likelihood of overfitting on small datasets. For high-quality fine-tuning on limited data, fewer targeted modules are usually preferred to maintain regularization and focus the training on the most impactful attention mechanisms within the model.
- D
Increase lora_alpha to 64
Why it fails: Increasing the alpha parameter effectively increases the weight of the LoRA adapters relative to the pre-trained weights. This makes the model more sensitive to the fine-tuning data, which would likely worsen overfitting rather than improving generalization. Alpha should generally be reduced or kept proportional to rank for stability.