MoDeGPT: Modular Decomposition for Large Language Model Compression
Chi-Heng Lin, Shangqian Gao, James Seale Smith, Abhishek Patel, Shikhar Tuli, Yilin Shen, Hongxia Jin, Yen-Chang Hsu
Abstract
Large Language Models (LLMs) have significantly advanced AI with their exceptional performance across a wide range of tasks. However, their extensive computational requirements restrict their use on devices with limited resources. While recent compression methods based on low-rank matrices show potential solutions, they often suffer from significant loss of accuracy or introduce substantial overhead in parameters and inference time. In this paper, we introduce Modular Decomposition (MoDeGPT), a new, efficient, and structured compression framework that overcomes these limitations. MoDeGPT jointly decomposes pairs of consecutive subcomponents within Transformer blocks, reduces hidden dimensions through output reconstruction on a larger structural scale than conventional low-rank methods, and repurposes three classical matrix decomposition algorithms-Nyström approximation, CR decomposition, and SVD-to ensure bounded errors in our novel decomposition approach. Our experiments show that MoDeGPT, without relying on backward propagation, consistently matches or surpasses the performance of prior techniques that depend on gradient information, while achieving a 98% reduction in compute costs when compressing a 13Bparameter model. On LLaMA-2/3 and OPT models, MoDeGPT retains 90-95% of zero-shot performance with compression rates of 25-30%. The compression process can be completed on a single GPU in a few hours, boosting inference throughput by up to 46%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f8230e9-77cd-4978-b0f2-db60c5ec3c7fCited by top-tier papers13
- Multi-Head Low-Rank AttentionSongtao Liu, Hongwu Peng, Zhiwei Zhang, Zhengyu Chen et al.ICLR 2026 · 18 citations
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin et al.ICLR 2026 · 6 citations
- Large Language Model Compression with Global Rank and Sparsity OptimizationChanghai Zhou, Qian Qiao, Yuhua Zhou, Yuxin Wu et al.ICLR 2026 · 6 citations
- The Curious Case of In-Training Compression of State Space ModelsMakram Chahine, Philipp Nazari, Daniela Rus, T. Konstantin RuschICLR 2026 · 4 citations
- Boomerang Distillation Enables Zero-Shot Model Size InterpolationSara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero et al.ICLR 2026 · 3 citations
Builds on24
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- A3: an Analytical Low-Rank Approximation Framework for AttentionJeffrey T. H. Wong, Cheng Zhang, Xinye Cao, Pedro Gimenes et al.ICML 2026 · 4 citations
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block SkippingYu-Chen Lu, Sheng-Feng Yu, Hui-Hsien Weng, Pei-Shuo Wang et al.AAAI 2026 · 1 citation
- Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM CompressionRuoling Qi, Yirui Liu, Xuaner Wu, Xiangyu Wang et al.ICML 2026
- Representation Drift Compensation: A Near-Zero Inference Cost Enhancement for LLM DecompositionXinhao Huang, You-Liang Huang, Zeyi WenICML 2026
- LatentLLM: Activation-Aware Transform to Multi-Head Latent AttentionToshiaki Koike-Akino, Xiangyu Chen, Jing Liu, Ye Wang et al.AAAI 2026 · 1 citation
