llama.cpp PR enables x86 VNNI for Q2_0; 8B decode throughput rises from 2.39 to 8.20 tok/s on Bonsai GGUFs
Read the original at old.reddit.com→I was going through the current llama.cpp CPU PRs and #26348 stood out because this isn't the usual +5% kernel optimization. It adds an x86 VNNI implementation for the Q2_0 × Q8_0 dot product, and the author's...
Original headline: "A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s"
Coverage timeline
- Aug 7, 12:27 UTC r/LocalLLaMA lead source A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s