Gemma 4 QAT handles KV cache quantization significantly better, according to KLD benchmarks comparing non-QAT and QAT configurations.
Read the original at old.reddit.com→Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlama.cpp v0.4.3, fork of llama.cpp with more KV cache quantization options, comparing Gemma Q4_0 non-QAT vs Gemma Q4_0...
Original headline: "Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show"
Coverage timeline
- Aug 12, 15:28 UTC r/LocalLLaMA lead source Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show