OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
Stephen Zhang, Vardan Papyan
Abstract
The recent paradigm shift to large-scale foundation models has brought about a new era for deep learning that, while has found great success in practice, has also been plagued by prohibitively expensive costs in terms of high memory consumption and compute. To mitigate these issues, there has been a concerted effort in post-hoc neural network pruning techniques that do not require costly retraining. Despite the considerable progress being made, existing methods often exhibit a steady drop in model performance as the compression increases. In this paper, we present a novel approach to compressing large transformers, coined OATS, that compresses the model weights by approximating each weight matrix as the sum of a sparse matrix and a low-rank matrix. Prior to the decomposition, the weights are first scaled by the second moment of their input embeddings, so as to ensure the preservation of outlier features recently observed in large transformer models. Without retraining, OATS achieves state-of-the-art performance when compressing large language models, such as Llama-3 and Phi-3, and vision transformers, such as Google's ViT and DINOv2, by up to 60%, all while speeding up the model's inference on a CPU by up to 1.37× compared to prior pruning methods. Our code is available at: https://github.com/stephenqz/OATS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Attention Sinks: A 'Catch, Tag, Release' Mechanism for EmbeddingsStephen Zhang, Mustafa Khan, Vardan PapyanNeurIPS 2025 · 18 citations
- Large Language Model Compression with Global Rank and Sparsity OptimizationChanghai Zhou, Qian Qiao, Yuhua Zhou, Yuxin Wu et al.ICLR 2026 · 6 citations
- A3: an Analytical Low-Rank Approximation Framework for AttentionJeffrey T. H. Wong, Cheng Zhang, Xinye Cao, Pedro Gimenes et al.ICML 2026 · 4 citations
- 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMsMehdi Makni, Xiang Meng, Rahul MazumderNeurIPS 2025 · 2 citations
- Compress Large Language Models via Collaboration Between Learning and Matrix ApproximationYuesen Liao, Zhiwei Li, Binrui Wu, Zihao Cheng et al.NeurIPS 2025 · 1 citation
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random GeneratorsRasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker et al.ICLR 2025
- NOLA: Compressing LoRA using Linear Combination of Random BasisSoroush Abbasi Koohpayegani, Navaneet K. L., Parsa Nooralinejad, Soheil Kolouri et al.ICLR 2024 · 33 citations
- MoDeGPT: Modular Decomposition for Large Language Model CompressionChi-Heng Lin, Shangqian Gao, James Seale Smith, Abhishek Patel et al.ICLR 2025
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 88 citations
