Qwen 3.8 Flash with next-GSQ-RCO-IQ2_XS achieves ~21 tok/s on RTX 3060 12GB + 16GB RAM (no gate pruning, 100% bit-exact)
Read the original at www.reddit.com→(Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!) About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could...
Original headline: "Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)"
Coverage timeline
- Oct 9, 18:41 UTC r/LocalLLaMA lead source Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)