SpecLA: Efficient speculative decoding for linear-attention models
Read the original at arxiv.org→arXiv:2607.16673v1 Announce Type: new Abstract: Linear-attention models replace the growing KV cache with recurrent states, but autoregressive decoding still reads, updates, and writes these states one token at a...
Original headline: "SpecLA: Efficient Speculative Decoding for Linear-Attention Models"