ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
Ashkan Shahbazi, Elaheh Akbari, Darian Salehi, Xinran Liu, Navid NaderiAlizadeh, Soheil Kolouri
Abstract
While self-attention has been instrumental in the success of Transformers, it can lead to overconcentration on a few tokens during training, resulting in suboptimal information flow. Enforcing doubly-stochastic constraints in attention matrices has been shown to improve structure and balance in attention distributions. However, existing methods rely on iterative Sinkhorn normalization, which is computationally costly. In this paper, we introduce a novel, fully parallelizable doublystochastic attention mechanism based on sliced optimal transport, leveraging Expected Sliced Transport Plans (ESP). Unlike prior approaches, our method enforces doubly stochasticity without iterative Sinkhorn normalization, significantly enhancing efficiency. To ensure differentiability, we incorporate a temperature-based soft sorting technique, enabling seamless integration into deep learning models. Experiments across multiple benchmark datasets, including image classification, point cloud classification, sentiment analysis, and neural machine translation, demonstrate that our enhanced attention regularization consistently improves performance across diverse applications. Our implementation code can be found at https: //github.com/dariansal/ESPFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Quantum Doubly Stochastic TransformersJannis Born, Filip Skogh, Kahn Rhrissorrakrai, Filippo Utro et al.NeurIPS 2025 · 7 citations
- FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPUFelix X.-F. Ye, Xingjie Li, An Yu, Ming-Ching Chang et al.ICML 2026 · 3 citations
- Mixed-Curvature Tree-Sliced Wasserstein DistanceDuy-Tung Pham, Viet-Hoang Tran, Thieu Vo, Tan NguyenICLR 2026
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- Expected Sliced Transport PlansXinran Liu, Rocio Diaz Martin, Yikun Bai, Ashkan Shahbazi et al.ICLR 2025
- Sparsity-Constrained Optimal TransportTianlin Liu, Joan Puigcerver, Mathieu BlondelICLR 2023 · 3 citations
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler et al.ICML 2020 · 391 citations
- Unlocking Slot Attention by Changing Optimal Transport CostsYan Zhang, David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts et al.ICML 2023 · 20 citations
- SeTformer Is What You Need for Vision and LanguagePourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael FelsbergAAAI 2024 · 8 citations
