WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control
Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen, Modi Shi, Yanjie Zhang, Delong Li, Chuanzhe Suo, wang chuang, zhihui peng, Hongyang Li
摘要
Humanoid robots require precise locomotion and dexterous manipulation to perform challenging locomanipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large-space loco-manipulation. We attribute this to: (1) the challenge of acquiring loco-manipulation knowledge due to the scarcity of humanoid teleoperation data, and (2) the difficulty of faithfully and reliably executing locomotion commands, stemming from the limited precision and stability of existing RL controllers. To acquire richer loco-manipulation knowledge, we propose a unified latent learning framework that enables Vision-Language-Action (VLA) system to learn from low-cost action-free egocentric videos. Moreover, an efficient data collection pipeline is devised to augment the dataset and scale the benefits. To more precisely execute the desired locomotion commands, we present a loco–manipulation–oriented (LMO) RL policy specifically tailored for accurate and stable core loco-manipulation movements, such as advancing, turning, and squatting. Building on these components, we introduce WholeBodyVLA, a unified framework for humanoid loco-manipulation. To the best of our knowledge, WholeBodyVLA is one of its kind enabling large-space humanoid loco–manipulation. It is verified via comprehensive experiments on the AgiBot X2 humanoid, outperforming prior baseline by 21.3%. It also demonstrates strong generalization and high extensibility across a broad range of tasks. Code and checkpoints would be made public.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder 等ICML 2024 · 被引用 513 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- Adversarial Locomotion and Motion Imitation for Humanoid Policy LearningJiyuan Shi, Xinzhe Liu, Dewei Wang, Ouyang Lu 等NeurIPS 2025 · 被引用 30 次
- Latent Action Pretraining from VideosSeonghyeon Ye, Joel Jang, Byeongguk Jeon, Se June Joo 等ICLR 2025
相关 Paper
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang 等CVPR 2026 · 被引用 14 次
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsXiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang 等ICLR 2026 · 被引用 59 次
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He 等CVPR 2026 · 被引用 14 次
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion RepresentationsShichao Fan, Kun Wu, Zhengping Che, Xinhua Wang 等ICML 2026 · 被引用 16 次
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang 等ICML 2026 · 被引用 3 次
