RBS-Attention: radius-bounded sparse prefill for long-context large language models
Read the original at arxiv.org→arXiv:2609.20971v1 Announce Type: new Abstract: Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse...
Original headline: "RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models"
Coverage timeline
- Sep 21, 04:00 UTC arXiv cs.AI lead source RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
- Sep 21, 04:00 UTC arXiv cs.LG Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding