Efficient Low-rank Backpropagation for Vision Transformer Adaptation
Yuedong Yang, Hung-Yueh Chiang, Guihong Li, Diana Marculescu, Radu Marculescu
Abstract
The increasing scale of vision transformers (ViT) has made the efficient finetuning of these large models for specific needs a significant challenge in various applications. This issue originates from the computationally demanding matrix multiplications required during the backpropagation process through linear layers in ViT. In this paper, we tackle this problem by proposing a new Low-rank Back-Propagation via Walsh-Hadamard Transformation (LBP-WHT) method. Intuitively, LBP-WHT projects the gradient into a low-rank space and carries out backpropagation. This approach substantially reduces the computation needed for adapting ViT, as matrix multiplication in the low-rank space is far less resource-intensive. We conduct extensive experiments with different models (ViT, hybrid convolution-ViT model) on multiple datasets to demonstrate the effectiveness of our method. For instance, when adapting an EfficientFormer-L1 model on CIFAR100, our LBP-WHT achieves 10.4% higher accuracy than the state-of-the-art baseline, while requiring 9 MFLOPs less computation. As the first work to accelerate ViT adaptation with low-rank backpropagation, our LBP-WHT method is complementary to many prior efforts and can be combined with them for better performance. How can we decrease the computational cost for all operations, including gradient computations for weights and inputs, involved in backpropagation (BP) through any linear layer in the ViT model? To answer this question, we introduce a new Low-rank BackPropagation via Walsh-Hadamard Transformation (LBP-WHT) method. As shown in Figure 1 , our method intuitively performs BP for gradients w.r.t. inputs and weights in a low-rank space. To achieve this, we project the gradient w.r.t. the output into a low-rank space using WHT [15], then perform low-rank matrix multiplications, and 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98dd8dac-0206-494d-b716-5e44e0f2da8bCited by top-tier papers8
- Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationQi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan et al.NeurIPS 2024 · 62 citations
- Activation Map Compression through Tensor Decomposition for Deep LearningLe-Trung Nguyen, Aël Quélennec, Enzo Tartaglione, Samuel Tardieu et al.NeurIPS 2024 · 7 citations
- Correlated Low-Rank Adaptation for ConvNetsWu Ran, Weijia Zhang, Shuyang Pang, Qi Zhu et al.NeurIPS 2025 · 5 citations
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski et al.ICLR 2026
- Efficient Resource-Constrained Training of Transformers via Subspace OptimizationLe-Trung Nguyen, Enzo Tartaglione, Van-Tam NguyenICLR 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyYiting Yang, Hao Luo, Yuan Sun, Qingsen Yan et al.ICCV 2025
- Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationWei Dong, Yuan Sun, Yiting Yang, Xing Zhang et al.NeurIPS 2024 · 10 citations
- Time-, Memory- and Parameter-Efficient Visual AdaptationOtniel-Bogdan Mercea, Alexey A. Gritsenko, Cordelia Schmid, Anurag ArnabCVPR 2024 · 11 citations
- Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesYuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu et al.ICLR 2025 · 10 citations
- Parameter-Efficient Model Adaptation for Vision TransformersXuehai He, Chunyuan Li, Pengchuan Zhang, Jianwei Yang et al.AAAI 2023 · 114 citations
