Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Jason Ramapuram, Federico Danieli, Eeshan Gunesh Dhekane, Floris Weers, Dan Busbridge, Pierre Ablin, Tatiana Likhomanenko, Jagrit Digani, Zijin Gu, Amitis Shidani, Russell Webb
摘要
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as the softmax of dot products between keys and queries. Recent work has explored alternatives to softmax attention in transformers, such as ReLU and sigmoid activations. In this work, we revisit sigmoid attention and conduct an in-depth theoretical and empirical analysis. Theoretically, we prove that transformers with sigmoid attention are universal function approximators and benefit from improved regularity compared to softmax attention. Through detailed empirical analysis, we identify stabilization of large initial attention norms during the early stages of training as a crucial factor for the successful training of models with sigmoid attention, outperforming prior attempts. We also introduce FLASH-SIGMOID, a hardware-aware and memory-efficient implementation of sigmoid attention yielding a 17% inference kernel speed-up over FLASHATTENTION2 on H100 GPUs 1 . Experiments across language, vision, and speech show that properly normalized sigmoid attention matches the strong performance of softmax attention on a wide range of domains and scales, which previous attempts at sigmoid attention were unable to fully achieve. Our work unifies prior art and establishes best practices for sigmoid attention as a drop-in softmax replacement in transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 被引用 85 次
- CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers UpSonghua Liu, Zhenxiong Tan, Xinchao WangNeurIPS 2025 · 被引用 33 次
- DuSA: Fast and Accurate Dual-Stage Sparse Attention Mechanism Accelerating Both Training and InferenceChong Wu, Jiawang Cao, Renjie Xu, Zhuoheng Ran 等NeurIPS 2025 · 被引用 5 次
- ZeroS: Zero-Sum Linear Attention for Efficient TransformersJiecheng Lu, Xu Han, Yan Sun, Viresh Pati 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Universal Approximation with Softmax AttentionJerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen, Weimin Wu 等ICML 2026 · 被引用 9 次
- Transformer Quality in Linear TimeWeizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. LeICML 2022 · 被引用 335 次
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured AttentionCan Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee 等NeurIPS 2025 · 被引用 5 次
- MetaAttention: A Unified and Performant Attention Framework across Hardware BackendsFeiyang Chen, Yu Cheng, Lei Wang, Yuqing Xia 等PPoPP 2026 · 被引用 1 次
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda 等ICML 2024 · 被引用 390 次
