Expected Sliced Transport Plans
Xinran Liu, Rocio Diaz Martin, Yikun Bai, Ashkan Shahbazi, Matthew Thorpe, Akram Aldroubi, Soheil Kolouri
Abstract
While self-attention has been instrumental in the success of Transformers, it can lead to overconcentration on a few tokens during training, resulting in suboptimal information flow. Enforcing doubly-stochastic constraints in attention matrices has been shown to improve structure and balance in attention distributions. However, existing methods rely on iterative Sinkhorn normalization, which is computationally costly. In this paper, we introduce a novel, fully parallelizable doublystochastic attention mechanism based on sliced optimal transport, leveraging Expected Sliced Transport Plans (ESP). Unlike prior approaches, our method enforces doubly stochasticity without iterative Sinkhorn normalization, significantly enhancing efficiency. To ensure differentiability, we incorporate a temperature-based soft sorting technique, enabling seamless integration into deep learning models. Experiments across multiple benchmark datasets, including image classification, point cloud classification, sentiment analysis, and neural machine translation, demonstrate that our enhanced attention regularization consistently improves performance across diverse applications. Our implementation code can be found at https: //github.com/dariansal/ESPFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- PLANETALIGN: A Comprehensive Python Library for Benchmarking Network AlignmentQi Yu, Zhichen Zeng, Yuchen Yan, Zhining Liu et al.ICLR 2026 · 12 citations
- Differentiable Generalized Sliced Wasserstein PlansLaetitia Chapel, Romain Tavenard, Samuel VaiterNeurIPS 2025 · 11 citations
- Fast Estimation of Wasserstein Distances via Regression on Sliced Wasserstein DistancesKhai Nguyen, Hai Nguyen, Nhat HoICLR 2026 · 5 citations
- ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport PlansAshkan Shahbazi, Elaheh Akbari, Darian Salehi, Xinran Liu et al.ICML 2025
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler et al.ICML 2020 · 391 citations
- Sparsity-Constrained Optimal TransportTianlin Liu, Joan Puigcerver, Mathieu BlondelICLR 2023 · 3 citations
- Unlocking Slot Attention by Changing Optimal Transport CostsYan Zhang, David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts et al.ICML 2023 · 20 citations
- FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPUFelix X.-F. Ye, Xingjie Li, An Yu, Ming-Ching Chang et al.ICML 2026 · 3 citations
- SeTformer Is What You Need for Vision and LanguagePourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael FelsbergAAAI 2024 · 8 citations
