On the Convergence of Stochastic Multi-Objective Gradient Manipulation and Beyond
Shiji Zhou, Wenpeng Zhang, Jiyan Jiang, Wenliang Zhong, Jinjie Gu, Wenwu Zhu
摘要
The conflicting gradients problem is one of the major bottlenecks for the effective training of machine learning models that deal with multiple objectives. To resolve this problem, various gradient manipulation techniques, such as PCGrad, MGDA, and CAGrad, have been developed, which directly alter the conflicting gradients to refined ones with alleviated or even no conflicts. However, the existing design and analysis of these techniques are mainly conducted under the full-batch gradient setting, ignoring the fact that they are primarily applied with stochastic mini-batch gradients. In this paper, we illustrate that the stochastic gradient manipulation algorithms may fail to converge to Pareto optimal solutions. Firstly, we show that these different algorithms can be summarized into a unified algorithmic framework, where the descent direction is given by the composition of the gradients of the multiple objectives. Then we provide an explicit two-objective convex optimization instance to explicate the non-convergence issue under the unified framework, which suggests that the non-convergence results from the determination of the composite weights solely by the instantaneous stochastic gradients. To fix the nonconvergence issue, we propose a novel composite weights determination scheme that exponentially averages the past calculated weights. Finally, we show the resulting new variant of stochastic gradient manipulation converges to Pareto optimal or critical solutions and yields comparable or improved empirical performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- FAMO: Fast Adaptive Multitask OptimizationBo Liu, Yihao Feng, Peter Stone, Qiang LiuNeurIPS 2023 · 被引用 127 次
- Revisiting Scalarization in Multi-Task Learning: A Theoretical PerspectiveYuzheng Hu, Ruicheng Xian, Qilong Wu, Qiuling Fan 等NeurIPS 2023 · 被引用 76 次
- Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-AvoidanceLisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi ChenNeurIPS 2023 · 被引用 53 次
- Direction-oriented Multi-objective Learning: Simple and Provable Stochastic AlgorithmsPeiyao Xiao, Hao Ban, Kaiyi JiNeurIPS 2023 · 被引用 46 次
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 被引用 41 次
它引用的顶会 Paper12
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual ModelsZirui Wang, Yulia Tsvetkov, Orhan Firat, Yuan CaoICLR 2021 · 被引用 241 次
- Towards Impartial Multi-task LearningLiyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue 等ICLR 2021 · 被引用 228 次
相关 Paper
- PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective OptimizationMingjing Xu, Peizhong Ju, Jia Liu, Haibo YangAAAI 2025 · 被引用 6 次
- Towards Task-Conflicts Momentum-Calibrated Approach for Multi-task LearningHeyan Chai, Zeyu Liu, Yongxin Tong, Ziyi Yao 等ICDE 2024 · 被引用 5 次
- Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent ApproachHeshan Devaka Fernando, Han Shen, Miao Liu, Subhajit Chaudhury 等ICLR 2023 · 被引用 2 次
- Multi-Objective Bilevel LearningZhiyao Zhang, Zhuqing Liu, Xin Zhang, Wen-Yen Chen 等AAAI 2026
- MGDA Converges under Generalized Smoothness, ProvablyQi Zhang, Peiyao Xiao, Shaofeng Zou, Kaiyi JiICLR 2025
