Tangent Transformers for Composition, Privacy and Removal
Tian Yu Liu, Aditya Golatkar, Stefano Soatto
Abstract
We introduce Tangent Attention Fine-Tuning (TAFT), a method for fine-tuning linearized transformers obtained by computing a First-order Taylor Expansion around a pre-trained initialization. We show that the Jacobian-Vector Product resulting from linearization can be computed efficiently in a single forward pass, reducing training and inference cost to the same order of magnitude as its original non-linear counterpart, while using the same number of parameters. Furthermore, we show that, when applied to various downstream visual classification tasks, the resulting Tangent Transformer fine-tuned with TAFT can perform comparably with fine-tuning the original non-linear network. Since Tangent Transformers are linear with respect to the new set of weights, and the resulting fine-tuning loss is convex, we show that TAFT enjoys several advantages compared to non-linear finetuning when it comes to model composition, parallel training, machine unlearning, and differential privacy. Our code is available at: https://github.com/ tianyu139/tangent-model-composition * Denotes equal contribution 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Tangent Model Composition for Ensembling and Continual Fine-tuningTian Yu Liu, Stefano SoattoICCV 2023 · 29 citations
- Dataless Weight Disentanglement in Task Arithmetic via Kronecker-Factored Approximate CurvatureAngelo Porrello, Pietro Buzzega, Felix Dangel, Thomas Sommariva et al.ICLR 2026 · 6 citations
- DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive SegmentationJihun Kim, Hoyong Kwon, Hyeokjun Kweon, Wooseong Jeong et al.ICCV 2025 · 1 citation
- Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust PlanningGiwon Lee, Wooseong Jeong, Daehee Park, Jaewoo Jeong et al.ICCV 2025 · 1 citation
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 1 citation
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
Related papers
- Distilling Linearized Behavior into Non-linear Fine-Tuning for Effective Task ArithmeticThomas Sommariva, Francesca Morandi, Simone Calderara, Angelo PorrelloICML 2026 · 2 citations
- Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task ArithmeticRuochen Jin, Bojian Hou, Jiancong Xiao, Weijie J. Su et al.ICLR 2025
- LQF: Linear Quadratic Fine-TuningAlessandro Achille, Aditya Golatkar, Avinash Ravichandran, Marzia Polito et al.CVPR 2021
- FLatten Transformer: Vision Transformer using Focused Linear AttentionDongchen Han, Xuran Pan, Yizeng Han, Shiji Song et al.ICCV 2023 · 358 citations
- A Second-Order Perspective on Model Compositionality and Incremental LearningAngelo Porrello, Lorenzo Bonicelli, Pietro Buzzega, Monica Millunzi et al.ICLR 2025
