SelKV: selective KV cache merging with per-token merge-or-drop and attention compensation
Read the original at arxiv.org→arXiv:2607.16213v1 Announce Type: new Abstract: Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a...
Original headline: "SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation"