Decode-Latency Feedback Prefill: A model-free controller and its generalization limits
Read the original at arxiv.org→arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already...
Original headline: "Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits"
Coverage timeline
- Oct 1, 04:00 UTC arXiv cs.AI lead source Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits