You Need Better Attention Priors
Elon Litman, Gabe Guo
摘要
We generalize the attention mechanism by viewing it through the lens of Entropic Optimal Transport, revealing that standard attention corresponds to a transport problem regularized by an implicit uniform prior. We introduce Generalized Optimal transport Attention with Trainable priors (GOAT ), a new attention mechanism that replaces this naive assumption with a learnable, continuous prior. This prior maintains full compatibility with optimized kernels such as FlashAttention. GOAT also provides an EOT-based explanation of attention sinks and materializes a solution for them, avoiding the representational trade-offs of standard attention. Finally, by absorbing spatial information into the core attention computation, GOAT learns an extrapolatable prior that combines the flexibility of learned positional embeddings with the length generalization of fixed encodings. arXiv preprint.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPUFelix X.-F. Ye, Xingjie Li, An Yu, Ming-Ching Chang 等ICML 2026 · 被引用 3 次
- MMReD: a Cross-Modal Benchmark for Dense Context ReasoningMaxim Kurkin, Boris Shirokikh, Irina Abdullaeva, Viktoriia Chekalina 等ICLR 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- Unlocking Slot Attention by Changing Optimal Transport CostsYan Zhang, David W. Zhang, Simon Lacoste-Julien, Gertjan J. Burghouts 等ICML 2023 · 被引用 20 次
- Accelerating 3D Molecule Generation via Jointly Geometric Optimal TransportHaokai Hong, Wanyu Lin, KC TanICLR 2025
- Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length ExtrapolationArthur S. Bianchessi, Yasmin C. Aguirre, Rodrigo C. Barros, Lucas S. KupssinsküICLR 2026 · 被引用 6 次
- Light Unbalanced Optimal TransportMilena Gazdieva, Arip Asadulaev, Evgeny Burnaev, Aleksandr KorotinNeurIPS 2024 · 被引用 9 次
- SCOPE: Boosting LLM Efficiency with Scoped Position EncodingQingguo Qi, Hongyang Chen, Zhao LiACL 2026
