HeraSys enables collaborative serving of multiple LLM workflows through fine-grained end-to-end optimization
Read the original at arxiv.org→arXiv:2607.22578v1 Announce Type: new Abstract: The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic...
Original headline: "HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization"
Coverage timeline
- Jul 28, 04:00 UTC arXiv cs.AI lead source HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization