Understanding Self-attention Mechanism via Dynamical System Perspective
Zhongzhan Huang, Mingfu Liang, Jinghui Qin, Shanshan Zhong, Liang Lin
摘要
The self-attention mechanism (SAM) is widely used in various fields of artificial intelligence and has successfully boosted the performance of different models. However, current explanations of this mechanism are mainly based on intuitions and experiences, while there still lacks direct modeling for how the SAM helps performance. To mitigate this issue, in this paper, based on the dynamical system perspective of the residual neural network, we first show that the intrinsic stiffness phenomenon (SP) in the high-precision solution of ordinary differential equations (ODEs) also widely exists in high-performance neural networks (NN). Thus the ability of NN to measure SP at the feature level is necessary to obtain high performance and is an important factor in the difficulty of training NN. Similar to the adaptive step-size method which is effective in solving stiff ODEs, we show that the SAM is also a stiffness-aware step size adaptor that can enhance the model's representational ability to measure intrinsic SP by refining the estimation of stiffness information and generating adaptive attention values, which provides a new understanding about why and how the SAM can benefit the model performance. This novel perspective can also explain the lottery ticket hypothesis in SAM, design new quantitative metrics of representational ability, and inspire a new theoretic-inspired approach, StepNet. Extensive experiments on several popular benchmarks demonstrate that StepNet can extract fine-grained stiffness information and measure SP accurately, leading to significant improvements in various visual tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local MinimaShanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen 等WWW 2024 · 被引用 24 次
- DVIB: Towards Robust Multimodal Recommender Systems via Variational Information Bottleneck DistillationWenkuan Zhao, Shanshan Zhong, Yifan Liu, Wushao Wen 等WWW 2025 · 被引用 7 次
- AttNS: Attention-Inspired Numerical Solving For Limited Data ScenariosZhongzhan Huang, Mingfu Liang, Shanshan Zhong, Liang LinICML 2024 · 被引用 6 次
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale ProblemsSaeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito KoishidaNeurIPS 2025 · 被引用 2 次
- Thermometer of Thoughts: Enhancing LLM's Exploration via Attention Temperature ModulationZhiyuan Yu, Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin 等ACL 2026
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si 等CVPR 2022 · 被引用 1,114 次
- FcaNet: Frequency Channel Attention NetworksZequn Qin, Pengyi Zhang, Fei Wu, Xi LiICCV 2021 · 被引用 1,049 次
相关 Paper
- Stiffness-aware neural network for learning Hamiltonian systemsSenwei Liang, Zhongzhan Huang, Hong ZhangICLR 2022 · 被引用 26 次
- Memory-Friendly Scalable Super-Resolution via Rewinding Lottery Ticket HypothesisJin Lin, Xiaotong Luo, Ming Hong, Yanyun Qu 等CVPR 2023
- Dynamical System Inspired Adaptive Time Stepping Controller for Residual Network FamiliesYibo Yang, Jianlong Wu, Hongyang Li, Xia Li 等AAAI 2020 · 被引用 23 次
- The Elastic Lottery Ticket HypothesisXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan 等NeurIPS 2021 · 被引用 38 次
- ResNet After All: Neural ODEs and Their Numerical SolutionKatharina Ott, Prateek Katiyar, Philipp Hennig, Michael TiemannICLR 2021 · 被引用 34 次
