Request-level energy attribution for batched LLM serving
Read the original at arxiv.org→arXiv:2608.00026v1 Announce Type: new Abstract: Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload...
Original headline: "Request-Level Energy Attribution for Batched LLM Serving"
Coverage timeline
- Aug 4, 04:00 UTC arXiv cs.AI lead source Request-Level Energy Attribution for Batched LLM Serving