Topology-aware data movement for disaggregated GPU inference
Read the original at arxiv.org→arXiv:2607.28633v1 Announce Type: new Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate GPU pools, the KV...
Original headline: "Topology-Aware Data Movement for Disaggregated GPU Inference"
Coverage timeline
- Aug 3, 04:00 UTC arXiv cs.LG lead source Topology-Aware Data Movement for Disaggregated GPU Inference