Can Small Heads Help? Understanding and Improving Multi-Task Generalization
Yuyan Wang, Zhe Zhao, Bo Dai, Christopher Fifty, Dong Lin, Lichan Hong, Li Wei, Ed H. Chi
Abstract
Multi-task learning aims to solve multiple machine learning tasks at the same time, with good solutions being both generalizable and Pareto optimal. A multi-task deep learning model consists of a shared representation learned to capture task commonalities, and task-specific sub-networks capturing the specificities of each task. In this work, we offer insights on the under-explored trade-off between minimizing task training conflicts in multi-task learning and improving multi-task generalization, i.e. the generalization capability of the shared presentation across all tasks. The trade-off can be viewed as the tension between multi-objective optimization and shared representation learning: As a multi-objective optimization problem, sufficient parameterization is needed for mitigating task conflicts in a constrained solution space; However, from a representation learning perspective, over-parameterizing the task-specific sub-networks may give the model too many ”degrees of freedom” and impedes the generalizability of the shared representation. Specifically, we first present insights on understanding the parameterization effect of multi-task deep learning models and empirically show that larger models are not necessarily better in terms of multi-task generalization. A delicate balance between mitigating task training conflicts vs. improving generalizability of the shared presentation learning is needed to achieve optimal performance across multiple tasks. Motivated by our findings, we then propose the use of a under-parameterized self-auxiliary head alongside each task-specific sub-network during training, which automatically balances the aforementioned trade-off. As the auxiliary heads are small in size and are discarded during inference time, the proposed method incurs minimal training cost and no additional serving cost. We conduct experiments with the proposed self-auxiliaries on two public datasets and live experiments on one of the largest industrial recommendation platforms serving billions of users. The results demonstrate the effectiveness of the proposed method in improving the predictive performance across multiple tasks in multi-task models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 593fcb66-3e54-4531-b1f5-af20f01e1523Cited by top-tier papers3
- Revisiting Scalarization in Multi-Task Learning: A Theoretical PerspectiveYuzheng Hu, Ruicheng Xian, Qilong Wu, Qiuling Fan et al.NeurIPS 2023 · 76 citations
- M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation FrameworkZijian Zhang, Shuchang Liu, Jiaao Yu, Qingpeng Cai et al.SIGIR 2024 · 27 citations
- Prompt and Parameter Co-Optimization for Large Language ModelsXiaohe Bo, Rui Li, Zexu Sun, Quanyu Dai et al.ICLR 2026 · 2 citations
Related papers
- Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto OptimizationDebabrata Mahapatra, Vaibhav RajanICML 2020 · 182 citations
- Learning Sparse Sharing Architectures for Multiple TasksTianxiang Sun, Yunfan Shao, Xiaonan Li, Pengfei Liu et al.AAAI 2020 · 155 citations
- Improving Gradient Trade-offs between Tasks in Multi-task Text ClassificationHeyan Chai, Jinhao Cui, Ye Wang, Min Zhang et al.ACL 2023 · 11 citations
- Learning with Privileged TasksYuru Song, Zan Lou, Shan You, Erkun Yang et al.ICCV 2021 · 3 citations
- MetaBalance: Improving Multi-Task Recommendations via Adapting Gradient Magnitudes of Auxiliary TasksYun He, Xue Feng, Cheng Cheng, Geng Ji et al.WWW 2022 · 69 citations
