Multi-level context modeling for consistent expert selection in Mixture-of-Experts
Read the original at arxiv.org→arXiv:2607.16427v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition...
Original headline: "Multi-level context Modeling for consistent expert selection in Mixture-of-Experts"