Beyond Sparse Weights: When Is Attention Compressible?
Read the original at arxiv.org→arXiv:2608.21541v1 Announce Type: new Abstract: KV-cache compression is often justified by attention maps with a few large weights. This is incomplete: large weights may not contain most of the mass, omitted values...
Coverage timeline
- Aug 25, 04:00 UTC arXiv cs.LG lead source Beyond Sparse Weights: When Is Attention Compressible?