Terminal shrinkage averaging reveals a schedule-estimator interaction in LLM pretraining
Read the original at arxiv.org→arXiv:2609.25482v1 Announce Type: new Abstract: Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the...
Original headline: "Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining"
Coverage timeline
- Sep 23, 04:00 UTC arXiv cs.LG lead source Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining