You Need Better Attention Priors
Elon Litman, Gabe Guo
Abstract
We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized by an implicit uniform prior. We introduce Generalized Optimal transport Attention with Trainable priors (GOAT ), a new attention mechanism that replaces this naive assumption with a learnable, continuous prior. This prior maintains full compatibility with optimized kernels such as FlashAttention. GOAT also provides an EOT-based explanation of attention sinks and materializes a solution for them, avoiding the representational trade-offs of standard attention. Finally, by absorbing spatial information into the core attention computation, GOAT learns an extrapolatable prior that combines the flexibility of learned positional embeddings with the length generalization of fixed encodings. arXiv preprint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04ae261a-3d5c-41a7-bbed-9e245302bb83Cited by top-tier papers2
- FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPUFelix X.-F. Ye, Xingjie Li, An Yu, Ming-Ching Chang et al.ICML 2026 · 3 citations
- MMReD: a Cross-Modal Benchmark for Dense Context ReasoningMaxim Kurkin, Boris Shirokikh, Irina Abdullaeva, Viktoriia Chekalina et al.ICLR 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Unlocking Slot Attention by Changing Optimal Transport CostsYan Zhang, David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts et al.ICML 2023 · 20 citations
- Accelerating 3D Molecule Generation via Jointly Geometric Optimal TransportHaokai Hong, Wanyu Lin, KC TanICLR 2025
- Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length ExtrapolationArthur S. Bianchessi, Yasmin C. Aguirre, Rodrigo C. Barros, Lucas S. KupssinsküICLR 2026 · 6 citations
- Light Unbalanced Optimal TransportMilena Gazdieva, Arip Asadulaev, Evgeny Burnaev, Aleksandr KorotinNeurIPS 2024 · 9 citations
- SCOPE: Boosting LLM Efficiency with Scoped Position EncodingQingguo Qi, Hongyang Chen, Zhao LiACL 2026
