WaveFront Decoding: a training-free self-speculative decoding framework for looped language models to reduce decoding latency by exploiting intermediate recurrence outputs.
Read the original at www.reddit.com→Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially...
Original headline: "[Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models"
Coverage timeline
- Oct 6, 13:18 UTC r/LocalLLaMA lead source [Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models