Keep On Going: Learning Robust Humanoid Motion Skills via Selective Adversarial Training
Yang Zhang, Zhanxiang Cao, Buqing Nie, Haoyang Li, Jiangwei Zhong, Qiao Sun, Xiaoyi Hu, Xiaokang Yang, Yue Gao
Abstract
Humanoid robots are expected to operate reliably over long horizons while executing versatile whole-body skills. Yet Reinforcement Learning (RL) motion policies typically lose stability under prolonged operation, sensor/actuator noise, and real world disturbances. In this work, we propose a Selective Adversarial Attack for Robust Training (SA2RT) to enhance the robustness of motion skills. The adversary is learned to identify and sparsely perturb the most vulnerable states and actions under an attack-budget constraint, thereby exposing true weakness without inducing conservative overfitting. The resulting non-zero sum, alternating optimization continually strengthens the motion policy against the strongest discovered attacks. We validate our approach on the Unitree G1 humanoid robot across perceptive locomotion and whole-body control tasks. Experimental results show that adversarially trained policies improve the terrain traversal success rate by 40%, reduce the trajectory tracking error by 32%, and maintain long horizon mobility and tracking performance. Together, these results demonstrate that selective adversarial attacks are an effective driver for learning robust, long horizon humanoid motion skills.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 161 citations
- Humanoid Locomotion as Next Token PredictionIlija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran et al.NeurIPS 2024 · 128 citations
- Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningYongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong HuangNeurIPS 2022 · 79 citations
- Adversarial Policy Learning in Two-player Competitive GamesWenbo Guo, Xian Wu, Sui Huang, Xinyu XingICML 2021 · 50 citations
- Adversarial Locomotion and Motion Imitation for Humanoid Policy LearningJiyuan Shi, Xinzhe Liu, Dewei Wang, Ouyang Lu et al.NeurIPS 2025 · 30 citations
Related papers
- Towards Robust and Safe Reinforcement Learning with Benign Off-policy DataZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICML 2023 · 14 citations
- Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning PolicyBuqing Nie, Yang Zhang, Rongjun Jin, Zhanxiang Cao et al.AAAI 2026 · 1 citation
- Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-TuningYingnan Zhao, Xinmiao Wang, Dewei Wang, Xinzhe Liu et al.AAAI 2026 · 4 citations
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICLR 2023 · 9 citations
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant et al.ICLR 2020 · 415 citations
