MGDA Converges under Generalized Smoothness, Provably
Qi Zhang, Peiyao Xiao, Shaofeng Zou, Kaiyi Ji
Abstract
Multi-objective optimization (MOO) is receiving more attention in various fields such as multi-task learning. Recent works provide some effective algorithms with theoretical analysis but they are limited by the standard L-smooth or boundedgradient assumptions, which typically do not hold for neural networks, such as Long short-term memory (LSTM) models and Transformers. In this paper, we study a more general and realistic class of generalized ℓ-smooth loss functions, where ℓ is a general non-decreasing function of gradient norm. We revisit and analyze the fundamental multiple gradient descent algorithm (MGDA) and its stochastic version with double sampling for solving the generalized ℓ-smooth MOO problems, which approximate the conflict-avoidant (CA) direction that maximizes the minimum improvement among objectives. We provide a comprehensive convergence analysis of these algorithms and show that they converge to an ϵ-accurate Pareto stationary point with a guaranteed ϵ-level average CA distance (i.e., the gap between the updating direction and the CA direction) over all iterations, where totally O(ϵ -2 ) and O(ϵ -4 ) samples are needed for deterministic and stochastic settings, respectively. We prove that they can also guarantee a tighter ϵ-level CA distance in each iteration using more samples. Moreover, we analyze an efficient variant of MGDA named MGDA-FA using only O(1) time and space, while achieving the same performance guarantee as MGDA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0de29556-6e4d-419a-bd25-b798bcf1fce0Cited by top-tier papers4
- Mirror Descent Under Generalized SmoothnessDingzhi Yu, Wei Jiang, Hongyi Tao, Yuanyu Wan et al.ICML 2026 · 9 citations
- Multi-objective Differentiable Neural Architecture SearchRhea Sanjay Sukthanker, Arber Zela, Benedikt Staffler, Samuel Dooley et al.ICLR 2025
- DREAM: A Unified Framework for Drift-Corrected Federated Multi-Objective LearningYuan Zhou, Yidan Ou, Xinli ShiICML 2026
- From Gradient Volume to Shapley Fairness: Towards Fair Multi-Task LearningXiao Wang, Yuying Han, Dazi Li, Fei Zhang et al.ICLR 2026
Builds on18
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong et al.NeurIPS 2020 · 313 citations
- Multi-Task Learning as a Bargaining GameAviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron et al.ICML 2022 · 243 citations
Related papers
- Direction-oriented Multi-objective Learning: Simple and Provable Stochastic AlgorithmsPeiyao Xiao, Hao Ban, Kaiyi JiNeurIPS 2023 · 46 citations
- Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-AvoidanceLisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi ChenNeurIPS 2023 · 53 citations
- Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent ApproachHeshan Devaka Fernando, Han Shen, Miao Liu, Subhajit Chaudhury et al.ICLR 2023 · 2 citations
- Adaptive Stochastic Gradient Algorithm for Black-box Multi-Objective LearningFeiyang Ye, Yueming Lyu, Xuehao Wang, Yu Zhang et al.ICLR 2024 · 5 citations
- Federated Multi-Objective LearningHaibo Yang, Zhuqing Liu, Jia Liu, Chaosheng Dong et al.NeurIPS 2023 · 28 citations
