Qwen 30b MoE runs at 30 tps on 6 GB VRAM with Hermes 90k context; user reports 20–35 tps during actual session generation
Read the original at old.reddit.com→So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of...
Original headline: "Qwen 30b MoE - 30tps - 6GB vram - Done!"
Coverage timeline
- Aug 14, 02:52 UTC r/LocalLLaMA lead source Qwen 30b MoE - 30tps - 6GB vram - Done!