CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters
Read the original at arxiv.org→arXiv:2610.00321v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters...
Coverage timeline
- Oct 2, 04:00 UTC arXiv cs.CL lead source CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters