ExFold: unified expert folding for training-free MoE prefill-decode acceleration
Read the original at arxiv.org→arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving...
Original headline: "ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration"
Coverage timeline
- Aug 27, 04:00 UTC arXiv cs.LG lead source ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration