Looped Latent Attention compresses K and V caches in looped transformers through a post-training codec, enabling compact cross-loop data storage
Read the original at arxiv.org→arXiv:2607.15456v1 Announce Type: new Abstract: Looped, weight-tied Transformers reduce parameters by reusing a block, but decoding still stores a separate K/V cache for every recurrence step. We show that this...
Original headline: "Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers"