GrowPage: on-demand KV budgeting for efficient LLM reasoning serving
Read the original at arxiv.org→arXiv:2609.03494v1 Announce Type: new Abstract: Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a...
Original headline: "GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.AI lead source GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving