Discretizing Continuous Action Space for On-Policy Optimization
Yunhao Tang, Shipra Agrawal
Abstract
In this work, we show that discretizing action space for continuous control is a simple yet powerful technique for on-policy optimization. The explosion in the number of discrete actions can be efficiently addressed by a policy with factorized distribution across action dimensions. We show that the discrete policy achieves significant performance gains with state-of-the-art on-policy optimization algorithms (PPO, TRPO, ACKTR) especially on high-dimensional tasks with complex dynamics. Additionally, we show that an ordinal parameterization of the discrete distribution can introduce the inductive bias that encodes the natural ordering between discrete actions. This ordinal architecture further significantly improves the performance of PPO/TRPO. An open source implementation of this paper can be found at https://github. com/robintyh1/onpolicybaselines .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a89cdb9a-84e1-495a-9628-bbc12894d057Cited by top-tier papers34
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain et al.ICML 2021 · 99 citations
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Learning Large Neighborhood Search Policy for Integer ProgrammingYaoxin Wu, Wen Song, Zhiguang Cao, Jie ZhangNeurIPS 2021 · 68 citations
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli PoliciesTim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato et al.NeurIPS 2021 · 59 citations
Related papers
- RN-D: Discretized Categorical Actors for On-Policy Reinforcement LearningYuexin Bian, Jie Feng, Tao Wang, Yijiang Li et al.ICML 2026
- Context-Sensitive Abstractions for Reinforcement Learning with Parameterized ActionsRashmeet Kaur Nayyar, Naman Shah, Siddharth SrivastavaAAAI 2026
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim et al.ICLR 2026 · 103 citations
- Subwords as Skills: Tokenization for Sparse-Reward Reinforcement LearningDavid Yunis, Justin Jung, Falcon Z. Dai, Matthew R. WalterNeurIPS 2024 · 5 citations
- Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action MaskingRoland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck et al.NeurIPS 2024 · 30 citations
