DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
Read the original at arxiv.org→arXiv:2608.05448v1 Announce Type: new Abstract: Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them....
Coverage timeline
- Aug 7, 04:00 UTC arXiv cs.CL lead source DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding