Keep On Going: Learning Robust Humanoid Motion Skills via Selective Adversarial Training
Yang Zhang, Zhanxiang Cao, Buqing Nie, Haoyang Li, Jiangwei Zhong, Qiao Sun, Xiaoyi Hu, Xiaokang Yang, Yue Gao
摘要
Humanoid robots are expected to operate reliably over long horizons while executing versatile whole-body skills. Yet Reinforcement Learning (RL) motion policies typically lose stability under prolonged operation, sensor/actuator noise, and real world disturbances. In this work, we propose a Selective Adversarial Attack for Robust Training (SA2RT) to enhance the robustness of motion skills. The adversary is learned to identify and sparsely perturb the most vulnerable states and actions under an attack-budget constraint, thereby exposing true weakness without inducing conservative overfitting. The resulting non-zero sum, alternating optimization continually strengthens the motion policy against the strongest discovered attacks. We validate our approach on the Unitree G1 humanoid robot across perceptive locomotion and whole-body control tasks. Experimental results show that adversarially trained policies improve the terrain traversal success rate by 40%, reduce the trajectory tracking error by 32%, and maintain long horizon mobility and tracking performance. Together, these results demonstrate that selective adversarial attacks are an effective driver for learning robust, long horizon humanoid motion skills.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 被引用 161 次
- Humanoid Locomotion as Next Token PredictionIlija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran 等NeurIPS 2024 · 被引用 128 次
- Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningYongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong HuangNeurIPS 2022 · 被引用 79 次
- Adversarial Policy Learning in Two-player Competitive GamesWenbo Guo, Xian Wu, Sui Huang, Xinyu XingICML 2021 · 被引用 50 次
- Adversarial Locomotion and Motion Imitation for Humanoid Policy LearningJiyuan Shi, Xinzhe Liu, Dewei Wang, Ouyang Lu 等NeurIPS 2025 · 被引用 30 次
相关 Paper
- Towards Robust and Safe Reinforcement Learning with Benign Off-policy DataZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICML 2023 · 被引用 14 次
- Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning PolicyBuqing Nie, Yang Zhang, Rongjun Jin, Zhanxiang Cao 等AAAI 2026 · 被引用 1 次
- Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-TuningYingnan Zhao, Xinmiao Wang, Dewei Wang, Xinzhe Liu 等AAAI 2026 · 被引用 4 次
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICLR 2023 · 被引用 9 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
