Calibrate, then route: a measured study of learned request routing for disaggregated LLM serving
Read the original at arxiv.org→arXiv:2609.16206v1 Announce Type: new Abstract: Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools. Systems such as DistServe, Splitwise, and Mooncake make this...
Original headline: "Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.AI lead source Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving