Vision-Language-Action Pretraining from Large-Scale Human Videos
Hao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng, Ye Wang, Haoqi Yuan, jiazheng liu, Chaoyi Xu, Haiweng Xu, Qin Jin, Zongqing Lu
Abstract
Existing Vision-Language-Action (VLA) models struggle with complex manipulation tasks requiring high dexterity and generalization, primarily due to their reliance on synthetic data with significant sim-to-real gaps or limited teleoperated demonstrations. To address this bottleneck, we propose leveraging human hands as a manipulator template, capitalizing on the rich dexterity and scalability present in web data of human manipulation. Our approach introduces physical instruction tuning, a novel training paradigm that combines large-scale VLA pretraining from human videos, perspective spatial alignment for reasoning in a unified physical space, and post-training adaptation in physical environments. Additionally, we introduce a part-level motion tokenization method that achieves millimeter-level reconstruction accuracy to model precise hand trajectories serving as scalable motion primitives. To support our paradigm, we develop a comprehensive data curation pipeline that integrates heterogeneous sources into a large-scale dataset with millions of motion-based instructional instances. Empirically, our model demonstrates superior performance in hand motion generation and instruction following, adhering to favorable scaling laws with respect to model and data sizes. Importantly, we demonstrate promising capabilities to robotic dexterous manipulation, validating the effectiveness of bridging the human-robot embodiment gap. Project page is available at https://research.beingbeyond.com/being-h0.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ff89fd2-1f1d-492a-9b7d-7d201619ae50Cited by top-tier papers5
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosGu Zhang, Qicheng Xu, Haozhe Zhang, Jianhan Ma et al.CVPR 2026 · 23 citations
- Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildHao Luo, Ye Wang, Wanpeng Zhang, Haoqi Yuan et al.CVPR 2026 · 15 citations
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human VideosYicheng Feng, Wanpeng Zhang, Ye Wang, Hao Luo et al.CVPR 2026 · 14 citations
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesXinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang et al.CVPR 2026 · 5 citations
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware RepresentationHaodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan et al.CVPR 2026 · 5 citations
Builds on67
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- UniHM: Unified Dexterous Hand Manipulation with Vision Language ModelZhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya WangICLR 2026 · 4 citations
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
- DexVLG: Dexterous Vision-Language-Grasp Model at ScaleJiawei He, Danshi Li, Xinqiang Yu, Zekun Qi et al.ICCV 2025 · 6 citations
- Cross-Hand Latent Representation for Vision-Language-Action ModelsGuangqi Jiang, Yutong Liang, Jianglong Ye, Jia-Yang Huang et al.CVPR 2026 · 14 citations
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong et al.ICLR 2026 · 20 citations
