Forgetfulness in fine-tuning: optimiser history causes loss of earlier answers due to stored gradients
Read the original at arxiv.org→arXiv:2610.03940v1 Announce Type: new Abstract: During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With...
Original headline: "A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.CL lead source A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass