Adaptive depth in looped transformers: diagnosing learned halting gates and trajectory readouts
Read the original at arxiv.org→arXiv:2607.20519v1 Announce Type: new Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block. Learned halting objectives in looped Transformers typically use a...
Original headline: "Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts"
Coverage timeline
- Jul 24, 04:00 UTC arXiv cs.LG lead source Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts