High-accuracy low-bit KV-cache quantization via local distribution restoration
Read the original at arxiv.org→arXiv:2607.16248v1 Announce Type: new Abstract: Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads. Low-bit...
Original headline: "High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration"
Coverage timeline
- Jul 21, 04:00 UTC arXiv cs.LG lead source High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration