Auxiliary Task Update Decomposition: the Good, the Bad and the neutral
Lucio M. Dery, Yann N. Dauphin, David Grangier
摘要
While deep learning has been very beneficial in data-rich settings, tasks with smaller training set often resort to pre-training or multitask learning to leverage data from other tasks. In this case, careful consideration is needed to select tasks and model parameterizations such that updates from the auxiliary tasks actually help the primary task. We seek to alleviate this burden by formulating a model-agnostic framework that performs fine-grained manipulation of the auxiliary task gradients. We propose to decompose auxiliary updates into directions which help, damage or leave the primary task loss unchanged. This allows weighting the update directions differently depending on their impact on the problem of interest. We present a novel and efficient algorithm for that purpose and show its advantage in practice. Our method leverages efficient automatic differentiation procedures and randomized singular value decomposition for scalability. We show that our framework is generic and encompasses some prior work as particular cases. Our approach consistently outperforms strong and widely used baselines when leveraging out-of-distribution data for Text and Image classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- ForkMerge: Mitigating Negative Transfer in Auxiliary-Task LearningJunguang Jiang, Baixu Chen, Junwei Pan, Ximei Wang 等NeurIPS 2023 · 被引用 55 次
- Should We Be Pre-training? An Argument for End-task Aware Training as an AlternativeLucio M. Dery, Paul Michel, Ameet Talwalkar, Graham NeubigICLR 2022 · 被引用 39 次
- Module-Aware Optimization for Auxiliary LearningHong Chen, Xin Wang, Yue Liu, Yuwei Zhou 等NeurIPS 2022 · 被引用 11 次
- Learning Conflict-Noticed Architecture for Multi-Task LearningZhixiong Yue, Yu Zhang, Jie LiangAAAI 2023 · 被引用 9 次
- A Bayesian Approach to Data Point SelectionXinnuo Xu, Minyoung Kim, Royson Lee, Brais Martínez 等NeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper5
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Unsupervised Pre-Training of Image Features on Non-Curated DataMathilde Caron, Piotr Bojanowski, Julien Mairal, Armand JoulinICCV 2019 · 被引用 254 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Optimizing Data Usage via Differentiable RewardsXinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos 等ICML 2020 · 被引用 73 次
- Learning a Multi-Domain Curriculum for Neural Machine TranslationWei Wang, Ye Tian, Jiquan Ngiam, Yinfei Yang 等ACL 2020 · 被引用 32 次
相关 Paper
- Auxiliary Learning by Implicit DifferentiationAviv Navon, Idan Achituve, Haggai Maron, Gal Chechik 等ICLR 2021 · 被引用 72 次
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property PredictionZhiqiang Zhong, Davide MottinNeurIPS 2025 · 被引用 3 次
- Learning with Privileged TasksYuru Song, Zan Lou, Shan You, Erkun Yang 等ICCV 2021 · 被引用 3 次
- MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary TasksYun He, Xue Feng, Cheng Cheng, Geng Ji 等WWW 2022 · 被引用 69 次
- Regularizing Neural Networks with Meta-Learning Generative ModelsShin'ya Yamaguchi, Daiki Chijiwa, Sekitoshi Kanai, Atsutoshi Kumagai 等NeurIPS 2023 · 被引用 10 次
