On the Convergence of Stochastic Multi-Objective Gradient Manipulation and Beyond
Shiji Zhou, Wenpeng Zhang, Jiyan Jiang, Wenliang Zhong, Jinjie Gu, Wenwu Zhu
Abstract
The conflicting gradients problem is one of the major bottlenecks for the effective training of machine learning models that deal with multiple objectives. To resolve this problem, various gradient manipulation techniques, such as PCGrad, MGDA, and CAGrad, have been developed, which directly alter the conflicting gradients to refined ones with alleviated or even no conflicts. However, the existing design and analysis of these techniques are mainly conducted under the full-batch gradient setting, ignoring the fact that they are primarily applied with stochastic mini-batch gradients. In this paper, we illustrate that the stochastic gradient manipulation algorithms may fail to converge to Pareto optimal solutions. Firstly, we show that these different algorithms can be summarized into a unified algorithmic framework, where the descent direction is given by the composition of the gradients of the multiple objectives. Then we provide an explicit two-objective convex optimization instance to explicate the non-convergence issue under the unified framework, which suggests that the non-convergence results from the determination of the composite weights solely by the instantaneous stochastic gradients. To fix the nonconvergence issue, we propose a novel composite weights determination scheme that exponentially averages the past calculated weights. Finally, we show the resulting new variant of stochastic gradient manipulation converges to Pareto optimal or critical solutions and yields comparable or improved empirical performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b572a12f-8be3-4118-895c-32c3be1dfb10Cited by top-tier papers24
- FAMO: Fast Adaptive Multitask OptimizationBo Liu, Yihao Feng, Peter Stone, Qiang LiuNeurIPS 2023 · 127 citations
- Revisiting Scalarization in Multi-Task Learning: A Theoretical PerspectiveYuzheng Hu, Ruicheng Xian, Qilong Wu, Qiuling Fan et al.NeurIPS 2023 · 76 citations
- Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-AvoidanceLisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi ChenNeurIPS 2023 · 53 citations
- Direction-oriented Multi-objective Learning: Simple and Provable Stochastic AlgorithmsPeiyao Xiao, Hao Ban, Kaiyi JiNeurIPS 2023 · 46 citations
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 41 citations
Builds on12
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 241 citations
- Towards Impartial Multi-task LearningLiyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue et al.ICLR 2021 · 228 citations
Related papers
- PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective OptimizationMingjing Xu, Peizhong Ju, Jia Liu, Haibo YangAAAI 2025 · 6 citations
- Towards Task-Conflicts Momentum-Calibrated Approach for Multi-task LearningHeyan Chai, Zeyu Liu, Yongxin Tong, Ziyi Yao et al.ICDE 2024 · 5 citations
- Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent ApproachHeshan Devaka Fernando, Han Shen, Miao Liu, Subhajit Chaudhury et al.ICLR 2023 · 2 citations
- Multi-Objective Bilevel LearningZhiyao Zhang, Zhuqing Liu, Xin Zhang, Wen-Yen Chen et al.AAAI 2026
- MGDA Converges under Generalized Smoothness, ProvablyQi Zhang, Peiyao Xiao, Shaofeng Zou, Kaiyi JiICLR 2025
