llamacpp runs slower than Ollama on a Huihui-Qwen3.6-35B-A3B-abliterated-ggml-model-Q4_K gguf setup, 55 tps vs. 61 tps for Ollama
Read the original at old.reddit.com→Hi. So I just setup llamacpp for the first time. I'm using the model : "Huihui-Qwen3.6-35B-A3B-abliterated-ggml-model-Q4_K.gguf". When I test this in llamacpp server GUI I get about 55tps, while in ollama default GUI...
Original headline: "llamacpp performing slower then Ollama"
Coverage timeline
- Aug 8, 16:02 UTC r/LocalLLaMA lead source llamacpp performing slower then Ollama