BeeLlama.cpp v0.4.1 adds KVarN, KV cache precision tail, and q2_0-q3_1/q6_0 cache support with improved benchmarks and VRAM efficiency
Read the original at old.reddit.com→TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more....
Original headline: "BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM"
Coverage timeline
- Jul 26, 16:28 UTC r/LocalLLaMA lead source BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM