Stiefel Attention: constraining transformer projection matrices to the Stiefel manifold with a Riemannian Adam update; four propositions show it yields steepest descent in the embedded metric regardless of gradient scale
Read the original at arxiv.org→arXiv:2609.19363v1 Announce Type: new Abstract: The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no constraint on their geometry. We constrain them to the...
Original headline: "Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not"
Coverage timeline
- Sep 18, 04:00 UTC arXiv cs.LG lead source Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not