Inferring DQN structure for high-dimensional continuous control
Andrey Sakryukin, Chedy Raïssi, Mohan S. Kankanhalli
Abstract
Despite recent advancements in the field of Deep Reinforcement Learning, Deep Q-network (DQN) models still show lackluster performance on problems with high-dimensional action spaces. The problem is even more pronounced for cases with high-dimensional continuous action spaces due to combinatorial increase in the number of the outputs. Recent works approach the problem by dividing the network into multiple parallel or sequential (action) modules responsible for different discretized actions. However there are drawbacks to both the parallel and the sequential approaches, i.e. parallel module architectures lack coordination between action modules, leading to extra complexity in the task, while a sequential structure can result in the vanishing gradients problem and exploding parameter space. In this work we show that the compositional structure of the action modules has a significant impact on the model performance, we propose a novel approach to infer the network structure for DQN models operating with high-dimensional continuous actions. Our method is based on uncertainty estimation techniques and yields substantially higher scores for MuJoCo environments with high-dimensional continuous action spaces, as well as a realistic AAA sailing simulator game.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 869f4d96-0a18-43d1-87f1-774b3cc929c5Cited by top-tier papers1
Ask how each one uses itRelated papers
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 25 citations
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 126 citations
- BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action SpacesMatthew Landers, Taylor W. Killian, Hugo Barnes, Tom Hartvigsen et al.NeurIPS 2025 · 6 citations
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 7 citations
- Simple Emergent Action Representations from Multi-Task Policy TrainingPu Hua, Yubei Chen, Huazhe XuICLR 2023 · 3 citations
