Auxiliary Task Update Decomposition: the Good, the Bad and the neutral
Lucio M. Dery, Yann N. Dauphin, David Grangier
Abstract
While deep learning has been very beneficial in data-rich settings, tasks with smaller training set often resort to pre-training or multitask learning to leverage data from other tasks. In this case, careful consideration is needed to select tasks and model parameterizations such that updates from the auxiliary tasks actually help the primary task. We seek to alleviate this burden by formulating a model-agnostic framework that performs fine-grained manipulation of the auxiliary task gradients. We propose to decompose auxiliary updates into directions which help, damage or leave the primary task loss unchanged. This allows weighting the update directions differently depending on their impact on the problem of interest. We present a novel and efficient algorithm for that purpose and show its advantage in practice. Our method leverages efficient automatic differentiation procedures and randomized singular value decomposition for scalability. We show that our framework is generic and encompasses some prior work as particular cases. Our approach consistently outperforms strong and widely used baselines when leveraging out-of-distribution data for Text and Image classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c9b842d-eff3-4d0b-9aa4-3a604a17d8c7Cited by top-tier papers10
- ForkMerge: Mitigating Negative Transfer in Auxiliary-Task LearningJunguang Jiang, Baixu Chen, Junwei Pan, Ximei Wang et al.NeurIPS 2023 · 55 citations
- Should We Be Pre-training? An Argument for End-task Aware Training as an AlternativeLucio M. Dery, Paul Michel, Ameet Talwalkar, Graham NeubigICLR 2022 · 39 citations
- Module-Aware Optimization for Auxiliary LearningHong Chen, Xin Wang, Yue Liu, Yuwei Zhou et al.NeurIPS 2022 · 11 citations
- Learning Conflict-Noticed Architecture for Multi-Task LearningZhixiong Yue, Yu Zhang, Jie LiangAAAI 2023 · 9 citations
- A Bayesian Approach to Data Point SelectionXinnuo Xu, Minyoung Kim, Royson Lee, Brais Martínez et al.NeurIPS 2024 · 3 citations
Builds on5
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Unsupervised Pre-Training of Image Features on Non-Curated DataMathilde Caron, Piotr Bojanowski, Julien Mairal, Armand JoulinICCV 2019 · 254 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Optimizing Data Usage via Differentiable RewardsXinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos et al.ICML 2020 · 73 citations
- Learning a Multi-Domain Curriculum for Neural Machine TranslationWei Wang, Ye Tian, Jiquan Ngiam, Yinfei Yang et al.ACL 2020 · 32 citations
Related papers
- Auxiliary Learning by Implicit DifferentiationAviv Navon, Idan Achituve, Haggai Maron, Gal Chechik et al.ICLR 2021 · 72 citations
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property PredictionZhiqiang Zhong, Davide MottinNeurIPS 2025 · 3 citations
- Learning with Privileged TasksYuru Song, Zan Lou, Shan You, Erkun Yang et al.ICCV 2021 · 3 citations
- MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary TasksYun He, Xue Feng, Cheng Cheng, Geng Ji et al.WWW 2022 · 69 citations
- Regularizing Neural Networks with Meta-Learning Generative ModelsShin'ya Yamaguchi, Daiki Chijiwa, Sekitoshi Kanai, Atsutoshi Kumagai et al.NeurIPS 2023 · 10 citations
