Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
Tairan He, Zi Wang, Haoru Xue, Qingwei Ben, Zhengyi Luo, Wenli Xiao, Ye Yuan, Xingye Da, Fernando Castañeda, Shankar Sastry, Changliu Liu, Guanya Shi
Abstract
A core barrier to the real-world productivity of humanoid robots is the lack of autonomous loco-manipulation skills. We introduce VIRAL, a visual sim-to-real framework that learns humanoid loco-manipulation entirely in simulation and deploys it zero-shot to real hardware. VIRAL follows a teacher-student design: a privileged RL teacher, operating on full state, learns long-horizon loco-manipulation using a delta action space and reference state initialization. A vision-based student policy is then distilled from the teacher via large-scale simulation with tiled rendering, trained with a mixture of online DAgger and behavior cloning. We find that compute scale is critical: scaling simulation to tens of GPUs (up to 64) makes both teacher and student training reliable, while low-compute regimes often fail. To bridge the sim-to-real gap, VIRAL combines large-scale visual domain randomization—over lighting, materials, camera parameters, image quality, and sensor delays—with real-to-sim alignment of the dexterous hand and cameras. Deployed on a Unitree G1 humanoid, the resulting RGB-based policy performs continuous loco-manipulation for up to 54 cycles, generalizing to diverse spatial and appearance variations without any real-world fine-tuning, and approaching expert-level teleoperation performance. Extensive ablations dissect the key design choices required to make RGB-based humanoid loco-manipulation work in practice.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ebf7fe12-063c-4f9c-828b-a5ef34348959Related papers
- Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy TransferHaoru Xue, Tairan He, Zi Wang, Qingwei Ben et al.CVPR 2026 · 35 citations
- GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and ImitationZifan Wang, Junyu Chen, Ziqing Chen, Pengwei Xie et al.CVPR 2024 · 15 citations
- Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement LearningPatrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino et al.ICLR 2026 · 15 citations
- WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation ControlHaoran Jiang, Jin Chen, Qingwen Bu, Li Chen et al.ICLR 2026 · 56 citations
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic SkillsJiayu Zhou, Qiwei Wu, Jian Li, Zhe Chen et al.AAAI 2026 · 1 citation
