Qwen 3.8 27B Q8 faster than Q6 with MTP on Apple Silicon using llama.cpp and lmstudio gguf
Read the original at old.reddit.com→Hi, Found something interesting while experimenting with Qwen 3.8 27B comparing Q8 and Q6 quant, using llama.cpp on 64GB Apple M1 Max. With MTP off, Q6 was faster than Q8 by about 10%, as expected. However with MTP...
Original headline: "Qwen 3.8 27B Q8 faster than Q6 w/ MTP on Apple Silicon using llama.cpp and lmstudio gguf"
Coverage timeline
- Aug 15, 21:14 UTC r/LocalLLaMA lead source Qwen 3.8 27B Q8 faster than Q6 w/ MTP on Apple Silicon using llama.cpp and lmstudio gguf