Poorman inference engine enables 35B Qwen 3.6 to run MoE expansion with 8 routed experts per token on a 16GB GPU; Claude Code compatibility via OpenAI and Anthropic APIs
Read the original at www.reddit.com→I forked llama.cpp's server into AgrillaMoE, a dedicated build for Qwen3.6-35B-A3B (~A4B) with Unsloth quants. On a (vant.ai) rented V100 16GB with the 2-bit UD-Q2_K_XL quant it generates at ~57-60 tok/s while...
Original headline: "poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding"
Coverage timeline
- Oct 4, 17:40 UTC r/LocalLLaMA lead source poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding