Lune

AAAI2020Top-tier venue

Discretizing Continuous Action Space for On-Policy Optimization

Yunhao Tang, Shipra Agrawal

2020Year
150Citations
34Top-tier citations

Abstract

In this work, we show that discretizing action space for continuous control is a simple yet powerful technique for on-policy optimization. The explosion in the number of discrete actions can be efficiently addressed by a policy with factorized distribution across action dimensions. We show that the discrete policy achieves significant performance gains with state-of-the-art on-policy optimization algorithms (PPO, TRPO, ACKTR) especially on high-dimensional tasks with complex dynamics. Additionally, we show that an ordinal parameterization of the discrete distribution can introduce the inductive bias that encodes the natural ordering between discrete actions. This ordinal architecture further significantly improves the performance of PPO/TRPO. An open source implementation of this paper can be found at https://github. com/robintyh1/onpolicybaselines .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a89cdb9a-84e1-495a-9628-bbc12894d057

Cited by top-tier papers34

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines