Continuous Control with Action Quantization from Demonstrations
Robert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin, Anton Raichuk, Matthieu Geist, Olivier Pietquin
Abstract
In this paper, we propose a novel Reinforcement Learning (RL) framework for problems with continuous action spaces: Action Quantization from Demonstrations (AQuaDem). The proposed approach consists in learning a discretization of continuous action spaces from human demonstrations. This discretization returns a set of plausible actions (in light of the demonstrations) for each input state, thus capturing the priors of the demonstrator and their multimodal behavior. By discretizing the action space, any discrete action deep RL technique can be readily applied to the continuous control problem. Experiments show that the proposed approach outperforms state-of-the-art methods such as SAC in the RL setup, and GAIL in the Imitation Learning setup. We provide a website with interactive videos: https://google-research.github.io/aquadem/ and make the code available: https://github.com/google-research/google-research/tree/master/aquadem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e416895-a98f-453a-9dac-e5bef47487dfCited by top-tier papers7
- Behavior Generation with Latent ActionsSeungjae Lee, Yibin Wang, Haritheja Etukuru, H. Jin Kim et al.ICML 2024 · 154 citations
- No Prior Mask: Eliminate Redundant Action for Deep Reinforcement LearningDianyu Zhong, Yiqin Yang, Qianchuan ZhaoAAAI 2024 · 15 citations
- Efficient Planning with Latent DiffusionWenhao LiICLR 2024 · 15 citations
- Investigating the Role of Model-Based Learning in Exploration and TransferJacob C. Walker, Eszter Vértes, Yazhe Li, Gabriel Dulac-Arnold et al.ICML 2023 · 8 citations
- Subwords as Skills: Tokenization for Sparse-Reward Reinforcement LearningDavid Yunis, Justin Jung, Falcon Z. Dai, Matthew R. WalterNeurIPS 2024 · 5 citations
Builds on15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu et al.ICLR 2021 · 161 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
Related papers
- Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-NetworkJijia Liu, Feng Gao, Qingmin Liao, Chao Yu et al.ICML 2025
- This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-CriticAndrea Marzo, Alessio Ragno, Roberto CapobiancoICML 2026
- Learning Human-Like RL Agents Through Trajectory Optimization With Action QuantizationJian-Ting Guo, Yu-Cheng Chen, Ping-Chun Hsieh, Kuo-Hao Ho et al.NeurIPS 2025 · 3 citations
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Learning Dialog Policies from Weak DemonstrationsGabriel Gordon-Hall, Philip John Gorinski, Shay B. CohenACL 2020 · 5 citations
