Expected Sliced Transport Plans
Xinran Liu, Rocio Diaz Martin, Yikun Bai, Ashkan Shahbazi, Matthew Thorpe, Akram Aldroubi, Soheil Kolouri
摘要
While self-attention has been instrumental in the success of Transformers, it can lead to overconcentration on a few tokens during training, resulting in suboptimal information flow. Enforcing doubly-stochastic constraints in attention matrices has been shown to improve structure and balance in attention distributions. However, existing methods rely on iterative Sinkhorn normalization, which is computationally costly. In this paper, we introduce a novel, fully parallelizable doublystochastic attention mechanism based on sliced optimal transport, leveraging Expected Sliced Transport Plans (ESP). Unlike prior approaches, our method enforces doubly stochasticity without iterative Sinkhorn normalization, significantly enhancing efficiency. To ensure differentiability, we incorporate a temperature-based soft sorting technique, enabling seamless integration into deep learning models. Experiments across multiple benchmark datasets, including image classification, point cloud classification, sentiment analysis, and neural machine translation, demonstrate that our enhanced attention regularization consistently improves performance across diverse applications. Our implementation code can be found at https: //github.com/dariansal/ESPFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PLANETALIGN: A Comprehensive Python Library for Benchmarking Network AlignmentQi Yu, Zhichen Zeng, Yuchen Yan, Zhining Liu 等ICLR 2026 · 被引用 12 次
- Differentiable Generalized Sliced Wasserstein PlansLaetitia Chapel, Romain Tavenard, Samuel VaiterNeurIPS 2025 · 被引用 11 次
- Fast Estimation of Wasserstein Distances via Regression on Sliced Wasserstein DistancesKhai Nguyen, Hai Nguyen, Nhat HoICLR 2026 · 被引用 5 次
- ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport PlansAshkan Shahbazi, Elaheh Akbari, Darian Salehi, Xinran Liu 等ICML 2025
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
相关 Paper
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler 等ICML 2020 · 被引用 391 次
- Sparsity-Constrained Optimal TransportTianlin Liu, Joan Puigcerver, Mathieu BlondelICLR 2023 · 被引用 3 次
- Unlocking Slot Attention by Changing Optimal Transport CostsYan Zhang, David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts 等ICML 2023 · 被引用 20 次
- FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPUFelix X.-F. Ye, Xingjie Li, An Yu, Ming-Ching Chang 等ICML 2026 · 被引用 3 次
- SeTformer Is What You Need for Vision and LanguagePourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael FelsbergAAAI 2024 · 被引用 8 次
