Attention-Aware Routing: Coupling routing and attention in MoEs
Read the original at arxiv.org→arXiv:2609.20974v1 Announce Type: new Abstract: In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We...
Original headline: "Attention-Aware Routing: Coupling Routing and Attention in MoEs"
Coverage timeline
- Sep 21, 04:00 UTC arXiv cs.AI lead source Attention-Aware Routing: Coupling Routing and Attention in MoEs