Proxyformer: Nyström-Based Linear Transformer with Trainable Proxy Tokens
Sangho Lee, Hayun Lee, Dongkun Shin
摘要
Transformer-based models have demonstrated remarkable performance in various domains, including natural language processing, image processing and generative modeling. The most significant contributor to the successful performance of Transformer models is the self-attention mechanism, which allows for a comprehensive understanding of the interactions between tokens in the input sequence. However, there is a well-known scalability issue, the quadratic dependency (i.e. O(n 2 )) of self-attention operations on the input sequence length n, making the handling of lengthy sequences challenging. To address this limitation, there has been a surge of research on efficient transformers, aiming to alleviate the quadratic dependency on the input sequence length. Among these, the Nyströmformer, which utilizes the Nyström method to decompose the attention matrix, achieves superior performance in both accuracy and throughput. However, its landmark selection exhibits redundancy, and the model incurs computational overhead when calculating the pseudo-inverse matrix. We propose a novel Nyström method-based transformer, called Proxyformer. Unlike the traditional approach of selecting landmarks from input tokens, the Proxyformer utilizes trainable neural memory, called proxy tokens, for landmarks. By integrating contrastive learning, input injection, and a specialized dropout for the decomposed matrix, Proxyformer achieves top-tier performance for long sequence tasks in the Long Range Arena benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning Generalizable Shape Completion with SIM(3) EquivarianceYuqing Wang, Zhaiyu Chen, Xiaoxiang ZhuNeurIPS 2025 · 被引用 2 次
- Unified Primitive Proxies for Structured Shape CompletionZhaiyu Chen, Yuqing Wang, Xiao Xiang ZhuCVPR 2026
它引用的顶会 Paper10
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
相关 Paper
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
- PoNet: Pooling Network for Efficient Token Mixing in Long SequencesChao-Hong Tan, Qian Chen, Wen Wang, Qinglin Zhang 等ICLR 2022 · 被引用 15 次
- ∞-former: Infinite Memory TransformerPedro Henrique Martins, Zita Marinho, André F. T. MartinsACL 2022 · 被引用 12 次
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström MethodYifan Chen, Qi Zeng, Heng Ji, Yun YangNeurIPS 2021 · 被引用 74 次
- Long-Short Transformer: Efficient Transformers for Language and VisionChen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi 等NeurIPS 2021 · 被引用 180 次
