HeadWiseKV: budgeted per-head cache residency for hybrid long-context language models
Read the original at arxiv.org→arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This...
Original headline: "HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models"
Coverage timeline
- Sep 3, 04:00 UTC arXiv cs.AI lead source HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models