Output-Aware Rotation for INT2 KV-Cache Quantization
Read the original at arxiv.org→arXiv:2608.02691v1 Announce Type: new Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization...
Coverage timeline
- Aug 5, 04:00 UTC arXiv cs.LG lead source Output-Aware Rotation for INT2 KV-Cache Quantization