SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
Jialong Guo, Xinghao Chen, Yehui Tang, Yunhe Wang
摘要
Transformers have become foundational architectures for both natural language and computer vision tasks. However, the high computational cost makes it quite challenging to deploy on resource-constraint devices. This paper investigates the computational bottleneck modules of efficient transformer, i.e., normalization layers and attention modules. LayerNorm is commonly used in transformer architectures but is not computational friendly due to statistic calculation during inference. However, replacing LayerNorm with more efficient BatchNorm in transformer often leads to inferior performance and collapse in training. To address this problem, we propose a novel method named PRepBN to progressively replace LayerNorm with re-parameterized BatchNorm in training. Moreover, we propose a simplified linear attention (SLA) module that is simple yet effective to achieve strong performance. Extensive experiments on image classification as well as object detection demonstrate the effectiveness of our proposed method. For example, our SLAB-Swin obtains top-1 accuracy on ImageNet-1K with ms latency, which is ms less than that of Flatten-Swin with higher accuracy. We also evaluated our method for language modeling task and obtain comparable performance and lower latency.Codes are publicly available at https://github.com/xinghaochen/SLAB and https://github.com/mindspore-lab/models/tree/master/research/huawei-noah/SLAB.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LiT: Delving into a Simple Linear Diffusion Transformer for Image GenerationJiahao Wang, Ning Kang, Lewei Yao, Mengzhao Chen 等ICCV 2025 · 被引用 10 次
- QT-ViT: Improving Linear Attention in ViT with Quadratic Taylor ExpansionYixing Xu, Chao Li, Dong Li, Xiao Sheng 等NeurIPS 2024 · 被引用 7 次
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- LaplacianFormer: Rethinking Linear Attention with Laplacian KernelZhe Feng, Sen Lian, Changwei Wang, Muyang Zhang 等ICLR 2026 · 被引用 3 次
- Vision Transformers Are Circulant Attention LearnersDongchen Han, Tianyu Li, Ziyi Wang, Gao HuangAAAI 2026 · 被引用 2 次
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
相关 Paper
- Unified Normalization for Accelerating and Stabilizing TransformersQiming Yang, Kai Zhang, Chaoxiang Lan, Zhi Yang 等ACM MM 2022 · 被引用 1 次
- Pre-RMSNorm and Pre-CRMSNorm Transformers: Equivalent and Efficient Pre-LN TransformersZixuan Jiang, Jiaqi Gu, Hanqing Zhu, David Z. PanNeurIPS 2023 · 被引用 44 次
- FLatten Transformer: Vision Transformer using Focused Linear AttentionDongchen Han, Xuran Pan, Yizeng Han, Shiji Song 等ICCV 2023 · 被引用 358 次
- Till the Layers Collapse: Compressing a Deep Neural Network Through the Lenses of Batch Normalization LayersZhu Liao, Nour Hezbri, Victor Quétu, Van-Tam Nguyen 等AAAI 2025 · 被引用 3 次
- ViDT: An Efficient and Effective Fully Transformer-based Object DetectorHwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani 等ICLR 2022 · 被引用 96 次
