QEvict: recoverable quantized KV eviction for attention-drift-robust long-context decoding
Read the original at arxiv.org→arXiv:2608.05326v1 Announce Type: new Abstract: Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this...
Original headline: "QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding"
Coverage timeline
- Aug 7, 04:00 UTC arXiv cs.LG lead source QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding