LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention
Toshiaki Koike-Akino, Xiangyu Chen, Jing Liu, Ye Wang, Pu Perry Wang, Matthew Brand
Abstract
Modern foundation models such as large language models (LLMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attention-aware joint tensor decomposition. Our framework can significantly improve the model accuracy over the existing model compression methods when reducing the latent dimension to realize computationally/memoryefficient LLMs. We show the benefit on several benchmark including multi-modal reasoning tasks.
- This work was done when X. Chen was an intern at MERL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 315 citations
Related papers
- Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model CompressionJingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li et al.ICLR 2025
- MoDeGPT: Modular Decomposition for Large Language Model CompressionChi-Heng Lin, Shangqian Gao, James Seale Smith, Abhishek Patel et al.ICLR 2025
- TD-MoE: Tensor Decomposition for MoE ModelsYuebin XU, YANHONG WANG, Xuemei Peng, Hui Zang et al.ICLR 2026
- ESPACE: Dimensionality Reduction of Activations for Model CompressionCharbel Sakr, Brucek KhailanyNeurIPS 2024 · 21 citations
- LeSTD: LLM Compression via Learning-based Sparse Tensor DecompositionYi Li, Zhichun Guo, Miao Yin, Bingzhe LiICLR 2026
