MoE pretraining reduces all-to-all overhead by correlating placement and token shuffling; expert parallelism cuts collective costs on multi-GPU clusters.
Read the original at arxiv.org→arXiv:2610.09372v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert...
Original headline: "Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling"
Coverage timeline
- Oct 8, 04:00 UTC arXiv cs.CL lead source Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling