TensorSharp adds MoE CPU-offload feature to main, enabling up to 35B-A3B MoE to fit with long-context KV cache on 12-16 GB GPUs; router, shared experts stay on accelerator.
Read the original at old.reddit.com→TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE expert weights of...
Original headline: "MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp"
Coverage timeline
- Aug 5, 13:13 UTC r/LocalLLaMA lead source MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp