REValueD: Regularised Ensemble Value-Decomposition for Factorisable Markov Decision Processes
David Ireland, Giovanni Montana
Abstract
Discrete-action reinforcement learning algorithms often falter in tasks with highdimensional discrete action spaces due to the vast number of possible actions. A recent advancement leverages value-decomposition, a concept from multi-agent reinforcement learning, to tackle this challenge. This study delves deep into the effects of this value-decomposition, revealing that whilst it curtails the overestimation bias inherent to Q-learning algorithms, it amplifies target variance. To counteract this, we present an ensemble of critics to mitigate target variance. Moreover, we introduce a regularisation loss that helps to mitigate the effects that exploratory actions in one dimension can have on the value of optimal actions in other dimensions. Our novel algorithm, REValueD, tested on discretised versions of the DeepMind Control Suite tasks, showcases superior performance, especially in the challenging humanoid and dog tasks. We further dissect the factors influencing REValueD's performance, evaluating the significance of the regularisation loss and the scalability of REValueD with increasing sub-actions per dimension. N i=1 n i atomic actions. Due to the combinatorial explosion of atomic actions that must be accounted for, standard algorithms such as Q-learning (Watkins and Dayan, 1992; Mnih et al., 2013) fail to learn in these settings as a result of computational impracticalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e4e9278-c4db-4b3a-8686-fba9a7b8a514Cited by top-tier papers4
- Stochastic Q-learning for Large Discrete Action SpacesFares Fourati, Vaneet Aggarwal, Mohamed-Slim AlouiniICML 2024 · 9 citations
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlChengxiu Hua, Jiawen Gu, Yushun TangNeurIPS 2025 · 5 citations
- Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement LearningMotoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta et al.ICML 2025
- Action-Free Offline-To-Online RL via Discretised State PoliciesNatinael Solomon Neggatu, Jeremie Houssineau, Giovanni MontanaICLR 2026
Builds on12
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li et al.NeurIPS 2020 · 142 citations
- Measuring the Reliability of Reinforcement Learning AlgorithmsStephanie C. Y. Chan, Samuel Fishman, Anoop Korattikara, John F. Canny et al.ICLR 2020 · 99 citations
Related papers
- Decoupling regularization from the action spaceSobhan Mohammadpour, Emma Frejinger, Pierre-Luc BaconICLR 2024 · 2 citations
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang et al.NeurIPS 2023 · 56 citations
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 9 citations
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 69 citations
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 56 citations
