Research · Updated 6 Oct, 09:30 am IST
Optimiser history causes loss of answer probability during fine‑tuning, study finds
Why it matters for readers: It explains why models sometimes ‘forget’ answers after fine‑tuning and that older optimizer states, not just current gradients, can cause this.
- Across three language‑model families, answer mass for old tasks consistently declines during fine‑tuning while discrimination among remaining answers often improves.1
- Decomposing Adam updates reveals opposing contributions: accumulated historical gradients favour leakage of probability outside the answer set, while the current gradient tends to oppose leakage.1
- Harmful contributions to forgetting come mainly from older gradients of the new task, whereas recent gradients tend to protect old answers.1
- Resetting momentum while matching the initial update norm reduces the harmful effect of stored history and improves final old‑task loss, largely by recovering answer mass.1
Get a brief like this every morning
Uzha reads hundreds of sources and gives you the stories that matter for your work, with every source linked. Free.
Get started