A3: an Analytical Low-Rank Approximation Framework for Attention
Jeffrey T. H. Wong, Cheng Zhang, Xinye Cao, Pedro Gimenes, Christos-Savvas Bouganis, George Constantinides, Wayne Luk, Aaron Zhao
摘要
Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-rank approximation offers a promising compression solution, yet existing approaches have two main limitations: (1) They focus on minimizing the output error of individual linear layers, without considering the architectural characteristics of Transformers, and (2) they decompose a large weight matrix into two small low-rank matrices. Consequently, these methods often fall short compared to other compression techniques like pruning and quantization, and introduce runtime overhead such as the extra GEMM kernel launches and memory operations for decomposed small matrices. To address these limitations, we propose , a post-training low-rank approximation framework. splits a Transformer layer into three functional components, namely , , and and provides analytical solutions that reduces the hidden dimension size inside each component while minimizing the component's functional loss. This approach directly reduces model sizes, KV cache sizes, and FLOPs without introducing any runtime overheads. Through extensive experiments, we show that maintains superior performance compared to SoTAs. For example, under the same reduction budget in computation and memory, our low-rank approximated LLaMA 3.1-70B achieves a perplexity of 4.69 on WikiText-2, outperforming the previous SoTA's 7.87 by 3.18. We also show versatile applications of in KV cache compression, integration with quantization, fine-tuning and mixed-rank assignments. We open-sourced our framework at https://github.com/DeepWok/a3.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model DeploymentRiccardo Zaccone, Stefanos Laskaridis, Marco Ciccone, Samuel HorváthICML 2026 · 被引用 1 次
- GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language ModelsZiyang Wang, Jiangfeng Xiao, Chuan Xiao, Ruoxiang Li 等ACL 2026
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou 等ICLR 2022 · 被引用 210 次
相关 Paper
- OATS: Outlier-Aware Pruning Through Sparse and Low Rank DecompositionStephen Zhang, Vardan PapyanICLR 2025
- 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMsMehdi Makni, Xiang Meng, Rahul MazumderNeurIPS 2025 · 被引用 2 次
- Palu: KV-Cache Compression with Low-Rank ProjectionChi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen 等ICLR 2025
- MoDeGPT: Modular Decomposition for Large Language Model CompressionChi-Heng Lin, Shangqian Gao, James Seale Smith, Abhishek Patel 等ICLR 2025
- Representation Drift Compensation: A Near-Zero Inference Cost Enhancement for LLM DecompositionXinhao Huang, You-Liang Huang, Zeyi WenICML 2026
