MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary Tasks
Yun He, Xue Feng, Cheng Cheng, Geng Ji, Yunsong Guo, James Caverlee
Abstract
In many personalized recommendation scenarios, the generalization ability of a target task can be improved via learning with additional auxiliary tasks alongside this target task on a multi-task network. However, this method often suffers from a serious optimization imbalance problem. On the one hand, one or more auxiliary tasks might have a larger influence than the target task and even dominate the network weights, resulting in worse recommendation accuracy for the target task. On the other hand, the influence of one or more auxiliary tasks might be too weak to assist the target task. More challenging is that this imbalance dynamically changes throughout the training process and varies across the parts of the same network. We propose a new method: MetaBalance to balance auxiliary losses via directly manipulating their gradients w.r.t the shared parameters in the multi-task network. Specifically, in each training iteration and adaptively for each part of the network, the gradient of an auxiliary loss is carefully reduced or enlarged to have a closer magnitude to the gradient of the target loss, preventing auxiliary tasks from being so strong that dominate the target task or too weak to help the target task. Moreover, the proximity between the gradient magnitudes can be flexibly adjusted to adapt MetaBalance to different scenarios. The experiments show that our proposed method achieves a significant improvement of 8.34% in terms of NDCG@10 upon the strongest baseline on two real-world datasets. The code of our approach can be found at here.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56fc201d-80d2-4fa6-a763-507d7a10f88cCited by top-tier papers15
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
- Multi-behavior Self-supervised Learning for RecommendationJingcao Xu, Chaokun Wang, Cheng Wu, Yang Song et al.SIGIR 2023 · 80 citations
- AdaTask: A Task-Aware Adaptive Learning Rate Approach to Multi-Task LearningEnneng Yang, Junwei Pan, Ximei Wang, Haibin Yu et al.AAAI 2023 · 70 citations
- Smooth Tchebycheff Scalarization for Multi-Objective OptimizationXi Lin, Xiaoyuan Zhang, Zhiyuan Yang, Fei Liu et al.ICML 2024 · 48 citations
Builds on3
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
- MTAdam: Automatic Balancing of Multiple Training Loss TermsItzik Malkiel, Lior WolfEMNLP 2021 · 12 citations
Related papers
- Towards Impartial Multi-task LearningLiyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue et al.ICLR 2021 · 228 citations
- Can Small Heads Help? Understanding and Improving Multi-Task GeneralizationYuyan Wang, Zhe Zhao, Bo Dai, Christopher Fifty et al.WWW 2022 · 15 citations
- Learning with Privileged TasksYuru Song, Zan Lou, Shan You, Erkun Yang et al.ICCV 2021 · 3 citations
- Auxiliary Learning as an Asymmetric Bargaining GameAviv Shamsian, Aviv Navon, Neta Glazer, Kenji Kawaguchi et al.ICML 2023 · 15 citations
- Quantifying Task Priority for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonCVPR 2024
