LeSTD: LLM Compression via Learning-based Sparse Tensor Decomposition
Yi Li, Zhichun Guo, Miao Yin, Bingzhe Li
Abstract
Large language models (LLMs) deliver the impressive capability, while their parameter scales hinder the deployment ability. Post-training matrix/tensor decomposition offers a promising strategy to alleviate this by exploiting structural redundancies within model weights. However, it faces the critical dense core bottleneck. This bottleneck caps achievable compression level as dense core tensor becomes a new storage burden. To solve this, we introduce LeSTD (Learningbased Sparse Tensor Decomposition), a two-stage, data-free compression framework. LeSTD first learns a high-quality shared basis for model weights, then applies a theoretically-ground pruning mechanism: guided by a derived closedform importance score, to create an ultra-sparse core tensor. Therefore, resulting a superior compression-accuracy trade-off: LeSTD achieves substantially higher compression ratios than dense-core methods without sacrificing performance. Experiments on the LLMs up to 30B parameters confirm that LeSTD consistently attains lower perplexity and higher task accuracy at matched compression levels, and critically, maintains strong performance under aggressive compression where prior methods degrade. Operationally, LeSTD executes the inference directly in its compressed domain, delivering significant throughput gains on standard hardware without requiring any custom kernels. along each mode, while the core tensor encodes the remaining cross-mode interactions. While the factor matrices can be relative small, the core tensor remains fully dense (Ahmadi-Asl et al., 2021; Wang & Yang, 2022) . To preserve model accuracy, the Tucker ranks must be sufficiently large, but the size of the core tensor grows polynomial with these ranks. This dense core rapidly becomes the new storage bottleneck, imposing a hard limit on the achievable compression ratio. This limitation exposes a fundamental gap: current tensor-based methods reduce dimensionality but fail to eliminate redundancy within the compressed latent space itself. To achieve the truly high-ratio compression, we must moving beyond low-rank approximation and into the domain of sparse tensor representation (Park et al., 2021) . To fill this critical gap, we propose LeSTD (Learning-based Sparse Tensor Decomposition), a framework that synergistically combines iterative basis optimization with the learned core tensor sparsity. LeSTD operates in two stages: first, it optimizes a high-quality shared basis for all attention heads; second, it learns an ultra-sparse representation for the core tensor within that basis. This integrated approach yields a representation that is both more accurate and vastly more compact. Main contributions of LeSTD are as follows: 1. The proposal of LeSTD, a data-free post-training compression framework where Stage I learns a high-quality shared basis via iterative optimization, and Stage II introduces a principled pruning strategy to create an ultra-sparse core tensor. 2. We provide a theoretical justification for the magnitude-based pruning in the Tucker-decomposed latent space. We derive a closed-form importance score for each core element, directly linking its magnitude to its impact on the Frobenius reconstruction error. This allows for a principled, rather than purely heuristic, sparsification of the core 3. We demonstrate how inference can operate directly on the compressed representation, avoiding the full weight reconstruction. It reduces arithmetic complexity and delivers practical throughput gains (tokens/sec) measured directly within the standard Transformers library (Wolf et al., 2020), requiring no specialized hardware or custom kernels. 4. Across GPT-J (6B), Llama2 (13B), and OPT (30B) on WikiText-2, MathQA, GSM8K, and Truth-fulQA, LeSTD consistently outperforms baselines at matched size fractions: maintaining higher accuracy under strong compression and delivering competitive-to-superior throughput. BACKGROUND AND PRELIMINARY Input Embedding Positional Encoding
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af204967-a248-4c2f-a91e-f1fb0dcde6fbBuilds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- MoDeGPT: Modular Decomposition for Large Language Model CompressionChi-Heng Lin, Shangqian Gao, James Seale Smith, Abhishek Patel et al.ICLR 2025
- SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model CompressionXinhao Huang, You-Liang Huang, Zeyi WenAAAI 2025 · 14 citations
- ESPACE: Dimensionality Reduction of Activations for Model CompressionCharbel Sakr, Brucek KhailanyNeurIPS 2024 · 21 citations
- LatentLLM: Activation-Aware Transform to Multi-Head Latent AttentionToshiaki Koike-Akino, Xiangyu Chen, Jing Liu, Ye Wang et al.AAAI 2026 · 1 citation
- CGSVD: Cascaded Granular Singular Value Decomposition for Large Language Model CompressionYuli Chen, Shuhao Zhang, Jiale Han, Fanshen Meng et al.ICML 2026
