From block-Toeplitz matrices to differential equations on graphs: towards a general theory for scalable masked Transformers
Krzysztof Choromanski, Han Lin, Haoxian Chen, Tianyi Zhang, Arijit Sehanobish, Valerii Likhosherstov, Jack Parker-Holder, Tamás Sarlós, Adrian Weller, Thomas Weingarten
摘要
In this paper we provide, to the best of our knowledge, the first comprehensive approach for incorporating various masking mechanisms into Transformers architectures in a scalable way. We show that recent results on linear causal attention (Choromanski et al., 2021) and log-linear RPE-attention (Luo et al., 2021) are special cases of this general mechanism. However by casting the problem as a topological (graph-based) modulation of unmasked attention, we obtain several results unknown before, including efficient d-dimensional RPE-masking and graph-kernel masking. We leverage many mathematical techniques ranging from spectral analysis through dynamic programming and random walks to new algorithms for solving Markov processes on graphs. We provide a corresponding empirical evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu 等NeurIPS 2022 · 被引用 1,216 次
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu 等NeurIPS 2023 · 被引用 97 次
- Polynormer: Polynomial-Expressive Graph Transformer in Linear TimeChenhui Deng, Zichao Yue, Zhiru ZhangICLR 2024 · 被引用 81 次
- Comparing Graph Transformers via Positional EncodingsMitchell Black, Zhengchao Wan, Gal Mishne, Amir Nayyeri 等ICML 2024 · 被引用 27 次
- On Oversquashing in Graph Neural Networks Through the Lens of Dynamical SystemsAlessio Gravina, Moshe Eliasof, Claudio Gallicchio, Davide Bacciu 等AAAI 2025 · 被引用 22 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
相关 Paper
- On the Emergence of Position Bias in TransformersXinyi Wu, Yifei Wang, Stefanie Jegelka, Ali JadbabaieICML 2025 · 被引用 1 次
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional EncodingShengjie Luo, Shanda Li, Tianle Cai, Di He 等NeurIPS 2021 · 被引用 66 次
- Linear Transformer Topological Masking with Graph Random FeaturesIsaac Reid, Kumar Avinava Dubey, Deepali Jain, William F. Whitney 等ICLR 2025
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing 等ICLR 2026 · 被引用 41 次
- StableMask: Refining Causal Masking in Decoder-only TransformerQingyu Yin, Xuzheng He, Xiang Zhuang, Yu Zhao 等ICML 2024 · 被引用 23 次
