Autonomy-of-Heads introduces data-free sparse attention from frozen query-key geometry to reduce long-context inference costs
Read the original at arxiv.org→arXiv:2608.06849v1 Announce Type: new Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods...
Original headline: "Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry"
Coverage timeline
- Aug 10, 04:00 UTC arXiv cs.CL lead source Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry