Adaptive batching helps LLM pretraining by allowing unbounded gradient variance, according to a perspective from unbounded variance
Read the original at arxiv.org→arXiv:2610.02355v1 Announce Type: new Abstract: Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not...
Original headline: "Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance"
Coverage timeline
- Oct 5, 04:00 UTC arXiv cs.LG lead source Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance