FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
Read the original at arxiv.org→arXiv:2610.06917v1 Announce Type: new Abstract: Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives....
Coverage timeline
- Oct 7, 04:00 UTC arXiv cs.AI lead source FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving