LeanStream: a speculate-and-refine streaming framework for efficient on-device LLM inference
Read the original at arxiv.org→arXiv:2609.03079v1 Announce Type: new Abstract: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available...
Original headline: "LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.LG lead source LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference