What Attention Recalls and Recurrence Controls in Hybrid Language Models; arXiv paper demonstrates two cache-level interventions (split-prefill and state-swap) on Qwen3.5 and Falcon-H1 to separate KV cache and recurrent state effects
Read the original at arxiv.org→arXiv:2609.04434v1 Announce Type: new Abstract: Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remains unclear. We introduce two cache-level interventions....
Original headline: "What Attention Recalls and Recurrence Controls in Hybrid Language Models"
Coverage timeline
- Sep 7, 04:00 UTC arXiv cs.CL lead source What Attention Recalls and Recurrence Controls in Hybrid Language Models