LinearKV: One cached state suffices for position-independent caching in hybrid LLMs
Read the original at arxiv.org→arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a...
Original headline: "LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs"
Coverage timeline
- Aug 13, 04:00 UTC arXiv cs.AI lead source LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs