FAMO: Fast Adaptive Multitask Optimization
Bo Liu, Yihao Feng, Peter Stone, Qiang Liu
Abstract
One of the grand enduring goals of AI is to create generalist agents that can learn multiple different tasks from diverse data via multitask learning (MTL). However, in practice, applying gradient descent (GD) on the average loss across all tasks may yield poor multitask performance due to severe under-optimization of certain tasks. Previous approaches that manipulate task gradients for a more balanced loss decrease require storing and computing all task gradients ( space and time where is the number of tasks), limiting their use in large-scale scenarios. In this work, we introduce Fast Adaptive Multitask Optimization FAMO, a dynamic weighting method that decreases task losses in a balanced way using space and time. We conduct an extensive set of experiments covering multi-task supervised and reinforcement learning problems. Our results indicate that FAMO achieves comparable or superior performance to state-of-the-art gradient manipulation techniques while offering significant improvements in space and computational efficiency. Code is available at https://github.com/Cranial-XIX/FAMO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7f3d9a8-a9c4-4a43-b4dd-21445e5fe76dCited by top-tier papers37
- GO4Align: Group Optimization for Multi-Task AlignmentJiayi Shen, Qi Wang, Zehao Xiao, Nanne van Noord et al.NeurIPS 2024 · 19 citations
- Neural Collapse in Multi-Task LearningYoujun Wang, Boqi Li, Xin Zou, Weiwei LiuICLR 2026 · 16 citations
- FERERO: A Flexible Framework for Preference-Guided Multi-Objective LearningLisha Chen, A F M Saif, Yanning Shen, Tianyi ChenNeurIPS 2024 · 12 citations
- PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient ConflictsZeman Li, Yuan Deng, Peilin Zhong, Meisam Razaviyayn et al.NeurIPS 2025 · 8 citations
- Efficient Utility-Preserving Machine Unlearning with Implicit Gradient SurgeryShiji Zhou, Tianbai Yu, Zhi Zhang, Heng Chang et al.NeurIPS 2025 · 6 citations
Builds on16
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas et al.ICML 2020 · 651 citations
- Efficiently Identifying Task Groupings for Multi-Task LearningChris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu et al.NeurIPS 2021 · 352 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
Related papers
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 41 citations
- Revisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance ReductionXiaohan Qin, Xiaoxing Wang, Junchi YanCVPR 2025
- Direction-oriented Multi-objective Learning: Simple and Provable Stochastic AlgorithmsPeiyao Xiao, Hao Ban, Kaiyi JiNeurIPS 2023 · 46 citations
- TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task OptimizationWonhyeok Choi, Kyumin Hwang, Jihun Park, Kyoungmin Lee et al.CVPR 2026
- AdaTask: A Task-Aware Adaptive Learning Rate Approach to Multi-Task LearningEnneng Yang, Junwei Pan, Ximei Wang, Haibin Yu et al.AAAI 2023 · 70 citations
