Margins, not windows: training-free per-step lossy speculative decoding
Read the original at arxiv.org→arXiv:2609.02897v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted,...
Original headline: "Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.CL lead source Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding