Qwen 3.6 27B Q5 runs on 3x2080ti with 55 tps using llama.cpp; user asks if more performance can be squeezed.
Read the original at old.reddit.com→CPU: Threadripper 3970X RAM: 128GB DDR4 GPUs: 3x2080ti 11GB The current best parameters to run it: llama-server \ --model Qwen3.6-27B-Q5_K_S.gguf \ --n-gpu-layers 999 \ --split-mode tensor \ --flash-attn on \...
Original headline: "Qwen 3.6 27B Q5 on 3x2080ti: 55tps with llama.cpp. Can I squeeze out more?"
Coverage timeline
- Aug 1, 07:52 UTC r/LocalLLaMA lead source Qwen 3.6 27B Q5 on 3x2080ti: 55tps with llama.cpp. Can I squeeze out more?