SOAP, Muon, and beyond: pushing LLM pretraining scales
Read the original at arxiv.org→arXiv:2607.20548v1 Announce Type: new Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited...
Original headline: "SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales"
Coverage timeline
- Jul 24, 04:00 UTC arXiv cs.LG lead source SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales