Meituan releases LongCat-Flash-Lite-Sparse MoE with ~3B active params and 30B n-gram lookup offloaded to RAM for fast 256k context on 24GB GPU
Read the original at old.reddit.com→It’s an MoE with ~3B active params and a 30B n-gram lookup table offloaded to RAM for fast 256k context on a 24GB GPU. Reminds me of Gemma 4’s PLE trick. Initial analysis suggest it wont be replacing my Qwen 3.6...
Original headline: "Meituan just dropped LongCat-Flash-Lite-Sparse"
Coverage timeline
- Jul 31, 14:46 UTC r/LocalLLaMA lead source Meituan just dropped LongCat-Flash-Lite-Sparse