Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies
Tim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato, Martin A. Riedmiller, Markus Wulfmeier, Daniela Rus
摘要
Reinforcement learning (RL) for continuous control typically employs distributions whose support covers the entire action space. In this work, we investigate the colloquially known phenomenon that trained agents often prefer actions at the boundaries of that space. We draw theoretical connections to the emergence of bang-bang behavior in optimal control, and provide extensive empirical evaluation across a variety of recent RL algorithms. We replace the normal Gaussian by a Bernoulli distribution that solely considers the extremes along each action dimension - a bang-bang controller. Surprisingly, this achieves state-of-the-art performance on several continuous control benchmarks - in contrast to robotic hardware, where energy and maintenance cost affect controller choices. Since exploration, learning,and the final solution are entangled in RL, we provide additional imitation learning experiments to reduce the impact of exploration on our analysis. Finally, we show that our observations generalize to environments that aim to model real-world challenges and evaluate factors to mitigate the emergence of bang-bang solutions. Our findings emphasize challenges for benchmarking continuous control algorithms, particularly in light of potential real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin 等ICML 2022 · 被引用 32 次
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay BuffersGautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He 等NeurIPS 2024 · 被引用 27 次
- Reinforcement Learning with Simple Sequence PriorsTankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz 等NeurIPS 2023 · 被引用 18 次
- Stochastic Q-learning for Large Discrete Action SpacesFares Fourati, Vaneet Aggarwal, Mohamed-Slim AlouiniICML 2024 · 被引用 9 次
- REValueD: Regularised Ensemble Value-Decomposition for Factorisable Markov Decision ProcessesDavid Ireland, Giovanni MontanaICLR 2024 · 被引用 6 次
它引用的顶会 Paper4
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 被引用 150 次
- Growing Action SpacesGregory Farquhar, Laura Gustafson, Zeming Lin, Shimon Whiteson 等ICML 2020 · 被引用 48 次
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 被引用 38 次
相关 Paper
- Truncated Gaussian Policy for Debiased Continuous ControlGanghun Lee, Minji Kim, Minsu Lee, Byoung-Tak ZhangAAAI 2025 · 被引用 1 次
- Analyzing Generalization in Policy Networks: A Case Study with the Double-Integrator SystemRuining Zhang, Haoran Han, Maolong Lv, Qisong Yang 等AAAI 2024 · 被引用 5 次
- What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale StudyMarcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini 等ICLR 2021 · 被引用 52 次
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
