Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
Jiaming Zhou, Teli Ma, Kun-Yu Lin, Zifan Wang, Ronghe Qiu, Junwei Liang
Abstract
lated tasks across two different benchmarks and five realworld tasks demonstrate significant improvements. These results span both single-task and language-conditioned multitask settings, evaluated using two different pre-trained models. Compared to existing pre-trained models, our adaptation method improves the average success rate by over 7% across multiple tasks on both simulated benchmarks and real-world evaluations. Project: https://jiamingzhou.github.io/projects/HumanRobotAlign
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c08e31b8-e07d-4a11-9907-9b6b4905ab34Cited by top-tier papers15
- Robotic Manipulation by Imitating Generated Videos Without Physical DemonstrationsShivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain et al.ICLR 2026 · 50 citations
- Improving Gloss-free Sign Language Translation by Reducing Representation DensityJinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang et al.NeurIPS 2024 · 49 citations
- Exploring the Limits of Vision-Language-Action Manipulation in Cross-task GeneralizationJiaming Zhou, Ke Ye, Jiayi Liu, Teli Ma et al.NeurIPS 2025 · 43 citations
- AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile PerceptionRuoxuan Feng, Yuxuan Zhou, Siyu Mei, Dongzhan Zhou et al.ICLR 2026 · 25 citations
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosGu Zhang, Qicheng Xu, Haozhe Zhang, Jianhan Ma et al.CVPR 2026 · 23 citations
Builds on19
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma et al.NeurIPS 2023 · 336 citations
- ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningJunting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao et al.NeurIPS 2022 · 290 citations
Related papers
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent GuidanceYang Zhang, Chenwei Wang, Ouyang Lu, Yuan Zhao et al.ICLR 2026 · 21 citations
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 1 citation
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure et al.ICLR 2024 · 114 citations
- LARA: Latent Action Representation Alignment for Vision-Language-Action ModelsMengya Liu, Baoxiong Jia, Jiangyong Huang, Jingze Zhang et al.ICML 2026 · 3 citations
- AnyBimanual: Transferring Unimanual Policy for General Bimanual ManipulationGuanxing Lu, Tengbo Yu, Haoyuan Deng, Season Si Chen et al.ICCV 2025 · 2 citations
