HybridInfer: Thermal-aware reinforcement-learning tier routing for on-device, edge, and cloud LLM inference
Read the original at arxiv.org→arXiv:2609.30270v1 Announce Type: new Abstract: On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is...
Original headline: "HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference"
Coverage timeline
- Sep 28, 04:00 UTC arXiv cs.LG lead source HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference