Quantizing KV Cache for DeepSeek V4 Flash degrades quality, per DS4F perplexity results
Read the original at old.reddit.com→I don't think anyone should quantize the KV with DS4F. I checked the the quality impact (PPL, KLD, Same TopP) for swhitching from BF16 KV to Q8 KV, and it appears significant. Very much in contrast to Qwen 397B. Here...
Original headline: "You really should not quantize KV Cache for DeepSeek V4 Flash"
Coverage timeline
- Aug 2, 22:01 UTC r/LocalLLaMA lead source You really should not quantize KV Cache for DeepSeek V4 Flash