Cacheable by design? Training mixture-of-experts routers for locality against the edge memory-bandwidth wall: a pre-registered negative result with a systems measurement study
Read the original at arxiv.org→arXiv:2608.18261v1 Announce Type: new Abstract: Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's...
Original headline: "Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study"