Qwen3.8-2.4T-A95B runs locally on RTX 5090 + RTX 5060 Ti at about 0.80 tokens per second using llama.cpp with Unsloth GGUF quantization and 512 routed experts, 10 active per token.
Read the original at old.reddit.com→I managed to get Qwen3.8-2.4T-A95B running locally with llama.cpp on mu PC just for fun, cause why not. I was using the Unsloth Qwen3.8-2.4T-A95B-UD-Q1_0 GGUF quantization. The full GGUF is about 397 GiB. The model...
Original headline: "EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s"
Coverage timeline
- Aug 13, 18:52 UTC r/LocalLLaMA lead source EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s