SpecLA: Efficient speculative decoding for linear-attention models
Read the original at arxiv.org→arXiv:2607.16673v1 Announce Type: new Abstract: Linear-attention models replace the growing KV cache with recurrent states, but autoregressive decoding still reads, updates, and writes these states one token at a...
Original headline: "SpecLA: Efficient Speculative Decoding for Linear-Attention Models"
Coverage timeline
- Jul 21, 04:00 UTC arXiv cs.CL lead source SpecLA: Efficient Speculative Decoding for Linear-Attention Models