DEP-RL: Embodied Exploration for Reinforcement Learning in Overactuated and Musculoskeletal Systems
Pierre Schumacher, Daniel F. B. Haeufle, Dieter Büchler, Syn Schmitt, Georg Martius
Abstract
Muscle-actuated organisms are capable of learning an unparalleled diversity of dexterous movements despite their vast amount of muscles. Reinforcement learning (RL) on large musculoskeletal models, however, has not been able to show similar performance. We conjecture that ineffective exploration in large overactuated action spaces is a key problem. This is supported by our finding that common exploration noise strategies are inadequate in synthetic examples of overactuated systems. We identify differential extrinsic plasticity (DEP), a method from the domain of self-organization, as being able to induce state-space covering exploration within seconds of interaction. By integrating DEP into RL, we achieve fast learning of reaching and locomotion in musculoskeletal systems, outperforming current approaches in all considered tasks in sample efficiency and robustness. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Latent exploration for Reinforcement LearningAlberto Silvio Chiappa, Alessandro Marin Vargas, Ann Zixiang Huang, Alexander MathisNeurIPS 2023 · 41 citations
- MyoDex: A Generalizable Prior for Dexterous ManipulationVittorio Caggiano, Sudeep Dasari, Vikash KumarICML 2023 · 26 citations
- DynSyn: Dynamical Synergistic Representation for Efficient Learning and Control in Overactuated Embodied SystemsKaibo He, Chenhui Zuo, Chengtian Ma, Yanan SuiICML 2024 · 19 citations
- Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement LearningGe Li, Hongyi Zhou, Dominik Roth, Serge Thilges et al.ICLR 2024 · 11 citations
- SIM2VR: Towards Automated Biomechanical Testing in VRFlorian Fischer, Aleksi Ikkala, Markus Klar, Arthur Fleig et al.UIST 2024 · 10 citations
Builds on4
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu et al.ICLR 2021 · 161 citations
- Growing Action SpacesGregory Farquhar, Laura Gustafson, Zeming Lin, Shimon Whiteson et al.ICML 2020 · 48 citations
- When should agents explore?Miruna Pislar, David Szepesvari, Georg Ostrovski, Diana L. Borsa et al.ICLR 2022 · 26 citations
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 25 citations
Related papers
- Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated ControlYiming Wang, Kaiyan Zhao, Xu Li, Yan Li et al.AAAI 2026 · 1 citation
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- Simplified Temporal Consistency Reinforcement LearningYi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala et al.ICML 2023 · 19 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
- Low-Rank Modular Reinforcement Learning via Muscle SynergyHeng Dong, Tonghan Wang, Jiayuan Liu, Chongjie ZhangNeurIPS 2022 · 21 citations
