Understanding energy scaling of large language model inference across context lengths and attention architectures
Read the original at arxiv.org→arXiv:2608.25096v1 Announce Type: new Abstract: The growing adoption of large language models (LLMs) has raised increasing concerns about the energy consumption and environmental impact of inference. This paper...
Original headline: "Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures"
Coverage timeline
- Aug 27, 04:00 UTC arXiv cs.LG lead source Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures