TreeSpark: Calibrated, load-adaptive draft trees for semi-autoregressive speculative decoding
Read the original at arxiv.org→arXiv:2609.22098v1 Announce Type: new Abstract: Speculative decoding accelerates language-model inference by letting a cheap drafter propose tokens that the target model verifies in parallel. Recent block drafters...
Original headline: "TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding"
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.CL lead source TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding