40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)
Read the original at old.reddit.com→daily reminder not to trust benchmarks and run it yourself. claimed e2e speedup is ~40%, forwards are ~140% faster I would wager that compared to a naive kernel anyone can write it's more in the range of 10-20%...
Coverage timeline
- Aug 5, 20:15 UTC r/LocalLLaMA lead source 40% speedup of MoE training with faster megakernel, by cursor, of all people (for B200s)