Faithful response-level interpretation of mixture-of-experts reward models via contribution contrast
Read the original at arxiv.org→arXiv:2608.06400v1 Announce Type: new Abstract: Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE)...
Original headline: "Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast"
Coverage timeline
- Aug 10, 04:00 UTC arXiv cs.AI lead source Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast