Running Qwen 3.5 35B A3B-Q8_0 gguf on a Radeon 7600 at 18 token/s with 64 GB DDR4 RAM and Ryzen 5600 using llama.cpp on Ubuntu
Read the original at old.reddit.com→I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 \ --n-cpu-moe 37 \ --no-mmap \ -ctk q8_0 \ -ctv q8_0 \ -fa 1 \ -c 9000 \ submitted by ...
Original headline: "Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s"
Coverage timeline
- Aug 10, 04:56 UTC r/LocalLLaMA lead source Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s