AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
Dejie Yang, Zijing Zhao, Yang Liu
摘要
Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multimodal data. To compensate for the deficiency of robot data, existing approaches have employed vision-language pretraining with large-scale data. However, they either utilize web data that differs from robotic tasks, or train the model in an implicit way (e.g., predicting future frames at the pixel level), thus showing limited generalization ability under insufficient robot data. In this paper, we propose to learn from large-scale human action video datasets in an explicit way (i.e., imitating human actions from hand keypoints), introducing Visual Robot Manipulation with Analogical Reasoning (AR-VRM). To acquire action knowledge explicitly from human action videos, we propose a keypoint Vision-Language Model (VLM) pretraining scheme, enabling the VLM to learn human action knowledge and directly predict human hand keypoints. During fine-tuning on robot data, to facilitate the robotic arm in imitating the action patterns of human motions, we first retrieve human action videos that perform similar manipulation tasks and have similar historical observations , and then learn the Analogical Reasoning (AR) map between human hand keypoints and robot components. Taking advantage of focusing on action keypoints instead of irrelevant visual cues, our method achieves leading performance on the CALVIN benchmark and real-world experiments. In few-shot scenarios, our AR-VRM outperforms previous methods by large margins , underscoring the effectiveness of explicitly imitating human actions under data scarcity. Code available at https://github.com/idejie/ar.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang 等CVPR 2026 · 被引用 6 次
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingMinghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng 等CVPR 2026 · 被引用 4 次
- Temporal-Aware Reasoning Optimization for Video Temporal GroundingMinghang Zheng, Zihao Yin, YI YANG, Yuxin Peng 等ICML 2026
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
相关 Paper
- VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained ActionsGuangyan Chen, Meiling Wang, Te Cui, Yao Mu 等NeurIPS 2024 · 被引用 24 次
- UP-VLA: A Unified Understanding and Prediction Model for Embodied AgentJianke Zhang, Yanjiang Guo, Yucheng Hu, Xiaoyu Chen 等ICML 2025
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji 等AAAI 2026 · 被引用 21 次
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosYi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li 等ICCV 2025 · 被引用 5 次
