Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
Haoru Xue, Tairan He, Zi Wang, Qingwei Ben, Wenli Xiao, Zhengyi Luo, Xingye Da, Fernando Castañeda, Guanya Shi, Shankar Sastry, Jim Fan, Yuke Zhu
Abstract
Recent progress in GPU-accelerated, photorealistic simulation has opened a scalable data-generation path for robot learning, where massive physics and visual randomization allow policies to generalize beyond curated environments. Building on these advances, we develop a teacher–student–bootstrap learning framework for vision-based humanoid loco-manipulation, using articulated-object interaction as a representative high-difficulty benchmark. Our approach introduces a staged-reset exploration strategy that stabilizes long-horizon privileged-policy training, and a GRPO-based fine-tuning procedure designed to mitigate partial observability and improve closed-loop consistency in sim-to-real RL. Trained entirely on synthetic simulation data, the resulting policy achieves robust zero-shot performance across diverse articulated objects—including multiple door types—and outperforms human teleoperators by up to 31.7% in task completion time under the same whole-body control stack. This represents the first humanoid sim-to-real policy capable of diverse articulated loco-manipulation from pure RGB perception.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccd631d4-4a03-4c8a-8696-d8703eb41d07Builds on2
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RLWenli Xiao, Haotian Lin, Andy Peng, Haoru Xue et al.ICLR 2026 · 84 citations
- Spectrum Random Masking for Generalization in Image-based Reinforcement LearningYangru Huang, Peixi Peng, Yifan Zhao, Guangyao Chen et al.NeurIPS 2022 · 33 citations
Related papers
- Visual Sim-to-Real at Scale for Humanoid Loco-ManipulationTairan He, Zi Wang, Haoru Xue, Qingwei Ben et al.CVPR 2026
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He et al.CVPR 2026 · 14 citations
- AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance CorrespondenceJiawei Zhang, Kaizhe Hu, Yingqian Huang, Yuanchen Ju et al.CVPR 2026
- HumanVLA: Towards Vision-Language Directed Object Rearrangement by Physical HumanoidXinyu Xu, Yizheng Zhang, Yonglu Li, Lei Han et al.NeurIPS 2024 · 29 citations
- Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement LearningPatrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino et al.ICLR 2026 · 15 citations
