When load-balancing goes too far: expert pruning fails in over-dispersed Mixture-of-Experts models
Read the original at arxiv.org→arXiv:2609.04453v1 Announce Type: new Abstract: Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router...
Original headline: "When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models"
Coverage timeline
- Sep 7, 04:00 UTC arXiv cs.LG lead source When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models