Compressing Transformers: Features Are Low-Rank, but Weights Are Not!
Hao Yu, Jianxin Wu
摘要
Transformer and its variants achieve excellent results in various computer vision and natural language processing tasks, but high computational costs and reliance on large training datasets restrict their deployment in resource-constrained settings. Low-rank approximation of model weights has been effective in compressing CNN models, but its application to transformers has been less explored and is less effective. Existing methods require the complete dataset to fine-tune compressed models, which are both time-consuming and data-hungry. This paper reveals that the features (i.e., activations) are low-rank, but model weights are surprisingly not low-rank. Hence, AAFM is proposed, which adaptively determines the compressed model structure and locally compresses each linear layer's output features rather than the model weights. A second stage, GFM, optimizes the entire compressed network holistically. Both AAFM and GFM only use few training samples without labels, that is, they are few-shot, unsupervised, fast and effective. For example, with only 2K images without labels, 33% of the parameters are removed in DeiT-B with 18.8% relative throughput increase, but only a 0.23% accuracy loss for ImageNet recognition. The proposed methods are successfully applied to the language modeling task in NLP, too. Besides, the few-shot compressed models generalize well in downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Small Singular Values Matter: A Random Matrix Analysis of Transformer ModelsMax Staats, Matthias Thamm, Bernd RosenowNeurIPS 2025 · 被引用 21 次
- LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language ModelsGuangyan Li, Yongqiang Tang, Wensheng ZhangICML 2024 · 被引用 11 次
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 被引用 10 次
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank StructureBoao Kong, Junzhu Liang, Yuxi Liu, Renjia Deng 等ICLR 2026 · 被引用 7 次
- ASER: Activation Smoothing and Error Reconstruction for Large Language Model QuantizationWeibo Zhao, Yubin Shi, Xinyu Lyu, Wanchen Sui 等AAAI 2025 · 被引用 7 次
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
相关 Paper
- Dense Vision Transformer Compression with Few SamplesHanxiao Zhang, Yifan Zhou, Guo-Hua WangCVPR 2024 · 被引用 5 次
- Parameter-Efficient Fine-Tuning with Discrete Fourier TransformZiqi Gao, Qichao Wang, Aochuan Chen, Zijing Liu 等ICML 2024 · 被引用 71 次
- Compressing Models with Few Samples: Mimicking then ReplacingHuanyu Wang, Junjie Liu, Xin Ma, Yang Yong 等CVPR 2022 · 被引用 11 次
- Variance-Based Pruning for Accelerating and Compressing Trained NetworksUranik Berisha, Jens Mehnert, Alexandru Paul ConduracheICCV 2025 · 被引用 7 次
- OATS: Outlier-Aware Pruning Through Sparse and Low Rank DecompositionStephen Zhang, Vardan PapyanICLR 2025
