Enhancing Transformers Through Conditioned Embedded Tokens
Hemanth Saratchandran, Simon Lucey
摘要
Transformers have transformed modern machine learning, driving breakthroughs in computer vision, natural language processing, and robotics. At the core of their success lies the attention mechanism, which enables the modeling of global dependencies among input tokens. However, we reveal that the attention block in transformers suffers from inherent ill-conditioning, which hampers gradient-based optimization and leads to inefficient training. To address this, we develop a theoretical framework that establishes a direct relationship between the conditioning of the attention block and that of the embedded tokenized data. Building on this insight, we introduce conditioned embedded tokens, a method that systematically modifies the embedded tokens to improve the conditioning of the attention mechanism. Our analysis demonstrates that this approach significantly mitigates ill-conditioning, leading to more stable and efficient training. We validate our methodology across various transformer architectures, achieving consistent improvements in image classification, object detection, instance segmentation, and natural language processing, highlighting its broad applicability and effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Spectral Conditioning of Attention Improves Transformer PerformanceHemanth Saratchandran, Simon LuceyNeurIPS 2025 · 被引用 9 次
- Always Skip AttentionYiping Ji, Hemanth Saratchandran, Peyman Moghadam, Simon LuceyICCV 2025
它引用的顶会 Paper17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski 等NeurIPS 2021 · 被引用 692 次
相关 Paper
- Conditioned Initialization for AttentionHemanth Saratchandran, Simon LuceyICLR 2026
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- Mitigating Over-smoothing in Transformers via Regularized Nonlocal FunctionalsTam Nguyen, Tan M. Nguyen, Richard G. BaraniukNeurIPS 2023 · 被引用 46 次
- Signal Propagation in Transformers: Theoretical Perspectives and the Role of Rank CollapseLorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto 等NeurIPS 2022 · 被引用 161 次
- Evolving Attention with Residual ConvolutionsYujing Wang, Yaming Yang, Jiangang Bai, Mingliang Zhang 等ICML 2021 · 被引用 43 次
