Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
Rui Dai, Sile Hu, Xu Shen, Yonggang Zhang, Xinmei Tian, Jieping Ye
Abstract
Task arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates that models demonstrating linearity enhance the performance of task arithmetic. In contrast to existing methods that rely on the global linearization of the model, we argue that this linearity already exists within the model's submodules. In particular, we present a statistical analysis and show that submodules (e.g., layers, self-attentions, and MLPs) exhibit significantly higher linearity than the overall model. Based on these findings, we propose an innovative model merging strategy that independently merges these submodules. Especially, we derive a closed-form solution for optimal merging weights grounded in the linear properties of these submodules. Experimental results demonstrate that our method consistently outperforms the standard task arithmetic approach and other established baselines across different model scales and various tasks. This result highlights the benefits of leveraging the linearity of submodules and provides a new perspective for exploring solutions for effective and practical multi-task model merging. * Corresponding authors. Code: https://github.com/deep-analysis-research/SLTA . 1 f (x; θ0 + ατ ) ≈ f (x; θ0) + α∆f (x; θ0 + τ ) where τ = θ -θ0 is the weight difference caused by fine-tuning and ∆f (x; θ0 + τ ) = f (x; θ0 + τ ) -f (x; θ0) is the corresponding output feature difference. 2 These two properties are easily confused. More discussions see Appendix A.5.1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- When Model Merging Breaks Routing: Training-Free Calibration for MoECanbin Huang, Tianyuan Shi, Xiaojun Quan, Jingang Wang et al.ICML 2026
- Scalable Model Merging with Progressive Layer-wise DistillationJing Xu, Jiazheng Li, Jingzhao ZhangICML 2025
- M-Loss: Quantifying Model Merging Compatibility with Limited Unlabeled DataTiantong Wang, Yiyang Duan, Haoyu Chen, Tiantong Wu et al.AAAI 2026
Builds on28
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task ArithmeticRuochen Jin, Bojian Hou, Jiancong Xiao, Weijie J. Su et al.ICLR 2025
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin et al.ICML 2024 · 96 citations
- MetaGPT: Merging Large Language Models Using Model Exclusive Task ArithmeticYuyan Zhou, Liang Song, Bingning Wang, Weipeng ChenEMNLP 2024 · 6 citations
- Revisiting the Role of Pretrained Weights in Model Merging: On Near-Optimality within the Core SubspaceWenju Sun, Qingyong Li, Tiancheng Li, Yangliao Geng et al.ICML 2026
