Fixed state, long reach: a constant-size cache enables block diffusion at scale
Read the original at arxiv.org→arXiv:2609.11998v1 Announce Type: new Abstract: Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the naive key--value (KV) cache behind fast autoregressive inference....
Original headline: "Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale"
Coverage timeline
- Sep 14, 04:00 UTC arXiv cs.LG lead source Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale