REValueD: Regularised Ensemble Value-Decomposition for Factorisable Markov Decision Processes
David Ireland, Giovanni Montana
摘要
Discrete-action reinforcement learning algorithms often falter in tasks with highdimensional discrete action spaces due to the vast number of possible actions. A recent advancement leverages value-decomposition, a concept from multi-agent reinforcement learning, to tackle this challenge. This study delves deep into the effects of this value-decomposition, revealing that whilst it curtails the overestimation bias inherent to Q-learning algorithms, it amplifies target variance. To counteract this, we present an ensemble of critics to mitigate target variance. Moreover, we introduce a regularisation loss that helps to mitigate the effects that exploratory actions in one dimension can have on the value of optimal actions in other dimensions. Our novel algorithm, REValueD, tested on discretised versions of the DeepMind Control Suite tasks, showcases superior performance, especially in the challenging humanoid and dog tasks. We further dissect the factors influencing REValueD's performance, evaluating the significance of the regularisation loss and the scalability of REValueD with increasing sub-actions per dimension. N i=1 n i atomic actions. Due to the combinatorial explosion of atomic actions that must be accounted for, standard algorithms such as Q-learning (Watkins and Dayan, 1992; Mnih et al., 2013) fail to learn in these settings as a result of computational impracticalities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Stochastic Q-learning for Large Discrete Action SpacesFares Fourati, Vaneet Aggarwal, Mohamed-Slim AlouiniICML 2024 · 被引用 9 次
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlChengxiu Hua, Jiawen Gu, Yushun TangNeurIPS 2025 · 被引用 5 次
- Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement LearningMotoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta 等ICML 2025
- Action-Free Offline-To-Online RL via Discretised State PoliciesNatinael Solomon Neggatu, Jeremie Houssineau, Giovanni MontanaICLR 2026
它引用的顶会 Paper12
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 被引用 150 次
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li 等NeurIPS 2020 · 被引用 142 次
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny 等ICLR 2020 · 被引用 99 次
相关 Paper
- Decoupling regularization from the action spaceSobhan Mohammadpour, Emma Frejinger, Pierre-Luc BaconICLR 2024 · 被引用 2 次
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang 等NeurIPS 2023 · 被引用 56 次
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 被引用 9 次
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 被引用 69 次
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 被引用 56 次
