Reformer: The Efficient Transformer
Nikita Kitaev, Lukasz Kaiser, Anselm Levskaya
摘要
Large Transformer models routinely achieve state-of-the-art results on a number of tasks but training these models can be prohibitively costly, especially on long sequences. We introduce two techniques to improve the efficiency of Transformers. For one, we replace dot-product attention by one that uses locality-sensitive hashing, changing its complexity from O() to O(), where is the length of the sequence. Furthermore, we use reversible residual layers instead of the standard residuals, which allows storing activations only once in the training process instead of times, where is the number of layers. The resulting model, the Reformer, performs on par with Transformer models while being much more memory-efficient and much faster on long sequences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper732
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 5,824 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
相关 Paper
- MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected LayersNing Ding, Yehui Tang, Haochen Qin, Zhenli Zhou 等NeurIPS 2024 · 被引用 8 次
- You Only Sample (Almost) Once: Linear Cost Self-Attention Via Bernoulli SamplingZhanpeng Zeng, Yunyang Xiong, Sathya N. Ravi, Shailesh Acharya 等ICML 2021 · 被引用 22 次
- SMYRF - Efficient Attention using Asymmetric ClusteringGiannis Daras, Nikita Kitaev, Augustus Odena, Alexandros G. DimakisNeurIPS 2020 · 被引用 54 次
- Treeformer: Dense Gradient Trees for Efficient Attention ComputationLovish Madaan, Srinadh Bhojanapalli, Himanshu Jain, Prateek JainICLR 2023
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
