NTK-approximating MLP Fusion for Efficient Language Model Fine-tuning
Tianxin Wei, Zeming Guo, Yifan Chen, Jingrui He
Abstract
Fine-tuning a pre-trained language model (PLM) emerges as the predominant strategy in many natural language processing applications. However, even fine-tuning the PLMs and doing inference are expensive, especially on edge devices with low computing power. Some general approaches (e.g. quantization and distillation) have been widely studied to reduce the compute/memory of PLM fine-tuning, while very few one-shot compression techniques are explored. In this paper, we investigate the neural tangent kernel (NTK)-which reveals the gradient descent dynamics of neural networks-of the multilayer perceptrons (MLP) modules in a PLM and propose to coin a lightweight PLM through NTKapproximating MLP fusion. To achieve this, we reconsider the MLP as a bundle of sub-MLPs, and cluster them into a given number of centroids, which can then be restored as a compressed MLP and surprisingly shown to well approximate the NTK of the original PLM. Extensive experiments of PLM fine-tuning on both natural language understanding (NLU) and generation (NLG) tasks are provided to verify the effectiveness of the proposed method MLP fusion. Our code is available at https://github.com/ weitianxin/MLP_Fusion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71ccd159-af5d-4354-9f27-939eb66fd813Cited by top-tier papers9
- Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondTianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng et al.ICLR 2024 · 46 citations
- Graph Mixup on Approximate Gromov-Wasserstein GeodesicsZhichen Zeng, Ruizhong Qiu, Zhe Xu, Zhining Liu et al.ICML 2024 · 30 citations
- SLOG: An Inductive Spectral Graph Neural Network Beyond Polynomial FilterHaobo Xu, Yuchen Yan, Dingsu Wang, Zhe Xu et al.ICML 2024 · 24 citations
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual RestorationMengting Ai, Tianxin Wei, Yifan Chen, Zhichen Zeng et al.KDD 2025 · 5 citations
- Scalable Multitask Learning Using Gradient-based Estimation of Task AffinityDongyue Li, Aneesh Sharma, Hongyang R. ZhangKDD 2024 · 3 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
Related papers
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
- Enabling Lightweight Fine-tuning for Pre-trained Language Model Compression based on Matrix Product OperatorsPeiyu Liu, Ze-Feng Gao, Wayne Xin Zhao, Zhi-Yuan Xie et al.ACL 2021
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
- Efficient Graph Continual Learning via Lightweight Graph Neural Tangent Kernels-based Dataset DistillationRihong Qiu, Xinke Jiang, Yuchen Fang, Hongbin Lai et al.ICML 2025
- NTK-SAP: Improving neural network pruning by aligning training dynamicsYite Wang, Dawei Li, Ruoyu SunICLR 2023 · 2 citations
