Minima-KV: retention-preserving KV cache compression with mixed-format paged attention
Read the original at arxiv.org→arXiv:2608.23834v1 Announce Type: new Abstract: The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We present Minima-KV, a retention-preserving hierarchy for...
Original headline: "Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention"
Coverage timeline
- Aug 26, 04:00 UTC arXiv cs.AI lead source Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
- Aug 26, 04:00 UTC arXiv cs.LG PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression