Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average Reward
Hairi, Jia Liu, Songtao Lu
摘要
Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multiobjective reinforcement learning (MORL) problem and introduces an innovative actor-critic algorithm named MOAC which finds a policy by iteratively making trade-offs among conflicting reward signals. Notably, we provide the first analysis of finite-time Pareto-stationary convergence and corresponding sample complexity in both discounted and average reward settings. Our approach has two salient features: (a) MOAC mitigates the cumulative estimation bias resulting from finding an optimal common gradient descent direction out of stochastic samples. This enables provable convergence rate and sample complexity guarantees independent of the number of objectives; (b) With proper momentum coefficient, MOAC initializes the weights of individual policy gradients using samples from the environment, instead of manual initialization. This enhances the practicality and robustness of our algorithm. Finally, experiments conducted on a real-world dataset validate the effectiveness of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Stochastic Linearized Augmented Lagrangian Method for Decentralized Bilevel OptimizationSongtao Lu, Siliang Zeng, Xiaodong Cui, Mark S. Squillante 等NeurIPS 2022 · 被引用 29 次
- Exploring both Individuality and Cooperation for Air-Ground Spatial Crowdsourcing by Multi-Agent Deep Reinforcement LearningYuxiao Ye, Chi Harold Liu, Zipeng Dai, Jianxin Zhao 等ICDE 2023 · 被引用 26 次
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu 等ICML 2024 · 被引用 4 次
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement LearningZhiyao Zhang, Myeung Suk Oh, Hairi, Ziyue Luo 等ICML 2025
它引用的顶会 Paper9
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 被引用 110 次
- On the Convergence of Stochastic Multi-Objective Gradient Manipulation and BeyondShiji Zhou, Wenpeng Zhang, Jiyan Jiang, Wenliang Zhong 等NeurIPS 2022 · 被引用 66 次
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue 等WWW 2023 · 被引用 60 次
- Taming Communication and Sample Complexities in Decentralized Policy Evaluation for Cooperative Multi-Agent Reinforcement LearningXin Zhang, Zhuqing Liu, Jia Liu, Zhengyuan Zhu 等NeurIPS 2021 · 被引用 36 次
- Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time AnalysisZiyi Chen, Yi Zhou, Rong-Rong Chen, Shaofeng ZouICML 2022 · 被引用 35 次
相关 Paper
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 被引用 8 次
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
- PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao LuoAAAI 2024 · 被引用 7 次
- Multi-Objective Reinforcement Learning with Max-Min Criterion: A Game-Theoretic ApproachWoohyeon Byeon, Giseung Park, Jongseong Chae, Amir Leshem 等NeurIPS 2025 · 被引用 6 次
