DAMP: decay-aware mixed-precision recurrent-state quantization
Read the original at arxiv.org→arXiv:2608.27513v1 Announce Type: new Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating...
Original headline: "DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization"
Coverage timeline
- Aug 31, 04:00 UTC arXiv cs.LG lead source DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
- Sep 1, 04:00 UTC arXiv cs.LG SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference