Lune

NeurIPS2025Top-tier venue

Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians

Akiyoshi Tomihari, Ryo Karakida

2025Year
5Citations
1Top-tier citations

Abstract

The theoretical understanding of self-attention (SA) has been steadily progressing. A prominent line of work studies a class of SA layers that admit an energy function decreased by state updates. While it provides valuable insights into inherent biases in signal propagation, it often relies on idealized assumptions or additional constraints not necessarily present in standard SA. Thus, to broaden our understanding, this work aims to relax these energy constraints and provide an energy-agnostic characterization of inference dynamics by dynamical systems analysis. In more detail, we first consider relaxing the symmetry and single-head constraints traditionally required in energy-based formulations. Next, we show that analyzing the Jacobian matrix of the state is highly valuable when investigating more general SA architectures without necessarily admitting an energy function. It reveals that the normalization layer plays an essential role in suppressing the Lipschitzness of SA and the Jacobian's complex eigenvalues, which correspond to the oscillatory components of the dynamics. In addition, the Lyapunov exponents computed from the Jacobians demonstrate that the normalized dynamics lie close to a critical state, and this criticality serves as a strong indicator of high inference performance. Furthermore, the Jacobian perspective also enables us to develop regularization methods for training and a pseudo-energy for monitoring inference dynamics.

is monotonically decreasing as dE multi (X)/dt ≤ 0 under the condition

where

Propositions 4.1 and 4.2 imply that certain structures of weight matrices are desirable to ensure the existence of an energy function. Specifically, W Q h W K⊤ h can be asymmetric, whereas W V h should remain symmetric. In the multi-head scenario, a low-rank structure in the QK product is required. This aligns with practical Transformers, as they typically exhibit a low-rank structure due to the small inner dimension (the width of W Q h , W K h ). We refer to architectures that incorporate these properties as generalized symmetric SA, and we will explore their effectiveness in our experiments (Section 6.2). The proofs are provided in Appendices A.2 and A.3.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f3e8f056-9cbf-4893-92b3-eca7d0953ae4

Cited by top-tier papers1

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines