Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy
Buqing Nie, Yang Zhang, Rongjun Jin, Zhanxiang Cao, Huangxuan Lin, Xiaokang Yang, Yue Gao
摘要
The human nervous system exhibits bilateral symmetry, enabling coordinated and balanced movements. However, existing Deep Reinforcement Learning (DRL) methods for humanoid robots neglect morphological symmetry of the robot, leading to uncoordinated and suboptimal behaviors. Inspired by human motor control, we propose Symmetry Equivariant Policy (SE-Policy), a new DRL framework that embeds strict symmetry equivariance in the actor and symmetry invariance in the critic without additional hyperparameters. SE-Policy enforces consistent behaviors across symmetric observations, producing temporally and spatially coordinated motions with higher task performance. Extensive experiments on velocity tracking tasks, conducted in both simulation and real-world deployment with the Unitree G1 humanoid robot, demonstrate that SE-Policy improves tracking accuracy by up to 40% compared to state-of-the-art baselines, while achieving superior spatial-temporal coordination. These results demonstrate the effectiveness of SE-Policy and its broad applicability to humanoid robots.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos 等AAAI 2021 · 被引用 506 次
- MDP Homomorphic Networks: Group Symmetries in Reinforcement LearningElise van der Pol, Daniel E. Worrall, Herke van Hoof, Frans A. Oliehoek 等NeurIPS 2020 · 被引用 203 次
- A Program to Build E(N)-Equivariant Steerable CNNsGabriele Cesa, Leon Lang, Maurice WeilerICLR 2022 · 被引用 133 次
- Continuous MDP Homomorphisms and Homomorphic Policy GradientSahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger 等NeurIPS 2022 · 被引用 34 次
- E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement LearningDingyang Chen, Qi ZhangICML 2024 · 被引用 10 次
相关 Paper
- Adversarial Locomotion and Motion Imitation for Humanoid Policy LearningJiyuan Shi, Xinzhe Liu, Dewei Wang, Ouyang Lu 等NeurIPS 2025 · 被引用 30 次
- KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic SkillsWeiji Xie, Jinrui Han, Jiakun Zheng, Huanyu Li 等NeurIPS 2025 · 被引用 120 次
- Hierarchical Equivariant Policy via Frame TransferHaibo Zhao, Dian Wang, Yizhe Zhu, Xupeng Zhu 等ICML 2025
- Keep On Going: Learning Robust Humanoid Motion Skills via Selective Adversarial TrainingYang Zhang, Zhanxiang Cao, Buqing Nie, Haoyang Li 等AAAI 2026 · 被引用 4 次
- ReActor: Reinforcement Learning for Physics-Aware Motion RetargetingDavid Müller, Agon Serifi, Sammy Christen, Ruben Grandia 等SIGGRAPH 2026
