Cursor releases their Mixture-of-Kittens megakernel for training MoE models; claims to nearly double TFLOP/s
Read the original at old.reddit.com→Link: https://cursor.com/blog/mixture-of-kittens GitHub: https://github.com/cursor/mixture-of-kittens Seems like a neat way to squeeze more performance out of MoE. I'm sure everyone has a favorite MoE model they'd...
Original headline: "Cursor releases their Mixture-of-Kittens megakernel for training MoE models - Claims to nearly double TFLOP/s"
Coverage timeline
- Aug 4, 17:26 UTC r/LocalLLaMA lead source Cursor releases their Mixture-of-Kittens megakernel for training MoE models - Claims to nearly double TFLOP/s