Mastering Task Arithmetic: τJp as a Key Indicator for Weight Disentanglement
Kotaro Yoshida, Yuji Naraki, Takafumi Horie, Ryosuke Yamaki, Ryotaro Shimizu, Yuki Saito, Julian J. McAuley, Hiroki Naganuma
Abstract
Model-editing techniques using task arithmetic have rapidly gained attention. Through task arithmetic, simply through arithmetic operations on the weights of pre-trained and fine-tuned models create desired models, such as multi-task models, models with specific tasks unsolvable, or domain-transferred models. However, task arithmetic faces challenges, such as low reproducibility and the high cost associated with adjusting coefficients in the arithmetic operations on model parameters, which have limited its practical success. In this paper, we present three key contributions in the context of task addition and task negation within task arithmetic. First, we propose a new metric called τ Jp which is based on the product of the task vector (τ ) and the Jacobian of the pre-trained model with respect to its weights. We show that τ Jp has a causal relationship with the interference that occurs from arithmetic operations. Second, we show that introducing regularization to minimize τ Jp significantly mitigates interference between task inferences, which leads to eliminating coefficient tuning and better accuracy on each task. Third, in the context of incremental learning, we confirmed that our τ Jp regularization demonstrates more robust performance in environments where future tasks to be learned are not accessible, validating the scalability of the approach. Finally, we demonstrate that the τ Jp regularizer further reinforces the performance of task arithmetic by leveraging publicly available fine-tuned models, offering practical benefits for real-world applications. Our code is available at https://github.com/katoro8989/tau-Jp_Task_Arithmetic * Work was performed when K.Yoshida and T.Horie were ProPlace interns 38th Workshop on Fine-Tuning in Machine Learning (NeurIPS 2024).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2a365d3-9632-4441-add2-698a289511e7Cited by top-tier papers5
- On Fairness of Task Arithmetic: The Role of Task VectorsLaura Gomezjurado Gonzalez, Hiroki Naganuma, Kotaro Yoshida, Takafumi Horie et al.ICLR 2026 · 3 citations
- Understanding and Enforcing Weight Disentanglement in Task ArithmeticShangge Liu, Yuehan Yin, Lei Wang, Qi Fan et al.CVPR 2026 · 3 citations
- ENCHTABLE: Unified Safety Alignment Transfer in Fine-Tuned Large Language ModelsJialin Wu, Kecen Li, Zhicong Huang, Xinfeng Li et al.S&P 2026 · 3 citations
- MergeRec: Model Merging for Data-Isolated Cross-Domain Sequential RecommendationHyunsoo Kim, Jaewan Moon, Seongmin Park, Jongwuk LeeKDD 2026
- Escaping Optimization Stagnation: Taking Steps Beyond Task Arithmetic via Difference VectorsJinping Wang, Zhiqiang Gao, Dinggen Zhang, Zhiwu XieAAAI 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear TransformersHongkang Li, Yihua Zhang, Shuai Zhang, Pin-Yu Chen et al.ICLR 2025
- Efficient Model Editing with Task-Localized Sparse Fine-tuningLeonardo Iurada, Marco Ciccone, Tatiana TommasiICLR 2025
- Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task ArithmeticRuochen Jin, Bojian Hou, Jiancong Xiao, Weijie J. Su et al.ICLR 2025
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
- Editing models with task arithmeticGabriel Ilharco, Marco Túlio Ribeiro, Mitchell Wortsman, Ludwig Schmidt et al.ICLR 2023 · 31 citations
